scylladb

Author	SHA1	Message	Date
Dani Tweig	5078a054d6	github: bug_report.yml: Improve bug report template structure V1: Perform a yaml "face lift" on the old bug report md template, making bug reporting more efficient. Add dedicated textarea fields for problem description and expected behavior Include pre-filled placeholders to guide issue reporting Add formatted log output section with shell syntax highlighting V2: updated the contact details of scylla and performed some code cleanup.	2024-11-14 12:47:06 +02:00
Dani Tweig	8cc510a51b	Update bug_report.yml Perform a yaml "face lift" to the old bug report md template. Asking to fill in addition to the former data details about reproduction steps and description of the problem. Making bug reporting more efficient.	2024-11-11 16:03:53 +02:00
Yaron Kaikov	cc71077e33	.github/scripts/label_promoted_commits.py: only match the Close tag in the last line in the commit message When a backport PR is promoted to the release branch, we automatically close the backport PR (since GitHub will only close the one based on the default branch) and update the labels in the original PRs In a situation when we have multiple `closes` prefixes, the script will use the first one (which is not the correct one), see `3ddb61c90e` Fixing this by always using the last line with the `closes` prefix Closes scylladb/scylladb#21498	2024-11-11 11:04:33 +02:00
Dani Tweig	381faa2649	Rename .github/ISSUE_TEMPLATE.md to .github/ISSUE_TEMPLATE/bug_report.yml GitHub issue template process has changed. The issue template file should be replaced and renamed. Closes scylladb/scylladb#21518	2024-11-11 11:00:38 +02:00
Nikita Kurashkin	3032d8ccbf	add check to refuse usage of DESC TABLE on a materialized view Fixes #21026 Closes scylladb/scylladb#21500	2024-11-11 10:23:30 +02:00
Yaron Kaikov	2596d1577b	./github/workflows/add-label-when-promoted.yaml: Run auto-backport only on default branch In https://github.com/scylladb/scylladb/pull/21496#event-15221789614 ``` scylladbbot force-pushed the backport/21459/to-6.1 branch from 414691c to `59a4ccd` Compare 2 days ago ``` Backport automation triggered by `push` but also should either start from `master` branch (or `enterprise` branch from Enterprise), we need to verify it by checking also the default branch. Fixes: https://github.com/scylladb/scylladb/issues/21514 Closes scylladb/scylladb#21515	2024-11-11 09:16:35 +02:00
Pavel Emelyanov	57af69e15f	Merge 'Add retries to the S3 client' from Ernest Zaslavsky 1. Add `retry_strategy` interface and default implementation for exponential back-off retry strategy. 2. Add new S3 related errors, also introduce additional errors to describe pure http errors that has no additional information in the body. 3. Add retries to the s3 client, all retries are coordinated by an instance of `retry_strategy`. In a case of error also parse response body in attempt to retrieve additional and more focused error information as suggested by AWS. See https://docs.aws.amazon.com/AmazonS3/latest/API/ErrorResponses.html. Introduce `aws_exception` to carry the original `aws_error`. 4. Discard whatever exception is thrown in `abort_upload` when aborting multipart upload since we don't care about cleanly aborting it since there are other means to clean up dangling parts, for example `rclone cleanup` or S3 bucket's Lifecycle Management Policy. 5. Add tests to cover retries, and retry exhaustion. Also add tests for jumbo upload. 6. Add the S3 proxy which is used to randomly inject retryable S3 errors to test the "retry" part of the S3 client. Switch the `s3_test` to use the S3 proxy. `s3_tests` set afloat `put_object` problem that was causing segmentation when retrying, fixed. 7. Extend the `s3_test` to use both `minio` and `proxy` configurations. 8. Add parameter to the proxy to seed the error injection randomization to make it replayable. fixes: #20611 fixes: #20613 Closes scylladb/scylladb#21054 * github.com:scylladb/scylladb: aws_errors: Make error messages more verbose. test: Make the minio proxy randomization re-playable test/boost/s3_test: add error injection scenarios to existing test suite test: Switch `s3_test` to use proxy test: Add more tests client: Stop returning error on `DELETE` in multipart upload abortion client: Fix sigsegv when retrying client: Add retries client: Adjust `map_s3_client_exception` to return exception instance aws_errors: Change aws_error::parse to return std::optional<> aws_errors: Add http errors mapping into aws_error client: Add aws_exception mapping aws_error: Add `aws_exeption` to carry original `aws_error` aws_errors: Add new error codes client: Introduce retry strategy	2024-11-11 08:35:55 +03:00
Takuya ASADA	92af373fab	unified: drop scylla-tools from unified package On `b8634fb`, we dropped scylla-tools from rpm and deb, we should drop it from unified package as well. Closes #20739 Closes scylladb/scylladb#20740	2024-11-10 12:56:43 +02:00
Avi Kivity	b58dbe57aa	Merge 'repair: introduce and use buffer size hint for mixed-shard multishard reader' from Botond Dénes Add a buffer hint to the multishard reader. This is an internal hint, used by the multishard reader to provide a hint to the shard reader, on how much data exactly is needed by the multishard reader from the respective shard. This hint allows eliminating extraneous cross-shard round-trips and possible shard reader evict-recreate cycles. Building on this, repair sets its own row buffer size as the max buffer size on the multishard reader, ensuring that the row buffer is filled with the minimum amount of cross-shard round trips and minimal reader recreation. To further eliminate unnecessary evictions, this PR also disables the multishard reader's read-ahead which is a mechanism that was designed to reduce latency for user-reads but it can be too aggressive for repair, causing unnecessary extra congestion on the already struggling streaming semaphores. Refs: https://github.com/scylladb/scylladb/issues/18269 Fixes: https://github.com/scylladb/scylladb/issues/21113 The performance impact was measured with an SCT test, which creates a cluster of 3 nodes with 16 shards, then adds a 4th one with 12 shards. Currently, it is the bootstrap time which is the worse in the case of mixed shard clusters, see below for the improvement measured during bootstrap: \| \| master \| buffer-hint \| metric \| \| ------------ \| ------------- \| ------------- \| --------------------------------------------------- \| \| evictions \| 0.9M \| 93.0K \| scylla_database_paused_reads_permit_based_evictions \| \| read (bytes) \| 9.0T \| 3.9T \| scylla_reactor_aio_bytes_read \| \| read (ops) \| 88.0M \| 33.5M \| scylla_reactor_aio_reads \| \| time \| 56min \| 20min \| N/A \| This is a performance improvement, no backport required. Closes scylladb/scylladb#20815 * github.com:scylladb/scylladb: test/boost/mutation_reader_test: add test for multishard reader buffer hint repair/row_level: disable read-ahead db/config: introduce repair_multishard_reader_enable_read_ahead readers/multishard: implement the read_ahead flag replica/database: make_multishard_streaming_reader(): expose the read_ahead parameter readers/multishard: add read_ahead parameter repair/row_level: set max buffer size on multishard reader replica/database: make_multishard_streaming_reader(): expose buffer_hint parameter db/config: introduce enable_repair_multishard_reader_buffer_hint readers/multishard: multishard_reader: pass hint to shard_reader readers/multishard: shard_reader_v2::fill_reader_buffer(): respect the hint readers/multishard: propagate fill_buffer_hint to shard_reader:fill_reader_buffer() readers/multishard: shard_reader: extract buffer-fill into its own method	2024-11-10 12:55:19 +02:00
Kefu Chai	961a53f716	dist: systemd: use default KillMode before this change, we specify the KillMode of the scylla-service service unit explicitly to "process". according to according to https://www.freedesktop.org/software/systemd/man/latest/systemd.kill.html, > If set to process, only the main process itself is killed (not recommended!). and the document suggests use "control-group" over "process". but scylla server is not a multi-process server, it is a multi-threaded server. so it should not make any difference even if we switch to the recommended "control-group". in the light that we've been seeing "defunct" scylla process after stopping the scylla service using systemd. we are wondering if we should try to change the `KillMode` to "control-group", which is the default value of this setting. in this change, we just drop the setting so that the systemd stops the service by stopping all processes in the control group of this unit are stopped. Refs scylladb/scylladb#21507 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21508	2024-11-09 20:07:11 +02:00
Kefu Chai	1f940d56b2	build: cmake: s/idle_compiler/idl_compiler/ before this change, the header files generated with `idl-compiler.py` are not regenerated if `idl-compiler.py` is updated. but they should, as the change to the script could in turn change the generated header files. because we have a typo in the `DEPENDS` argument, `${idle_compiler}` is expanded to an empty string. in this change, the typo is corrected, and the dependency from the generated headers to the script is correctly reflected in the building rules. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21475	2024-11-09 20:06:23 +02:00
Piotr Dulikowski	7021efd6b0	Merge 'main,cql_test_env: start group0_service before view_builder' from Michał Jadwiszczak In scylladb/scylladb#19745, view_builder was migrated to group0 and since then it is dependant on group0_service. Because of this, group0_service should be initialized/destroyed before/after view_builder. This patch also adds error injection to `raft_server_with_timeouts::read_barrier`, which does 1s sleep before doing the read barrier. There is a new test which reproduces the use after free bug using the error injection. Fixes scylladb/scylladb#20772 scylladb/scylladb#19745 is present in 6.2, so this fix should be backported to it. Closes scylladb/scylladb#21471 * github.com:scylladb/scylladb: test/boost/secondary_index_test: add test for use after free api/raft: use `get_server_with_timeouts().read_barrier()` in coroutines main,cql_test_env: start group0_service before view_builder	2024-11-08 20:27:09 +01:00
Kefu Chai	aebb532906	bytes, utils: include fmt/iostream.h and iostream when appropriate in seastar e96932b05f394b27cd0101e24f0584736795b50f, we stopped including unused `fmt/ostream.h`. this helped to reduce the header dependency. but this also broke the build of scylladb, as we rely on the `fmt/ostream.h` indirectly included by seastar's header project. in this change, we include `fmt/iostream.h` and `iostream` explictly when we are using the declarations in them. this enables us to - bump up the seastar submodule - potentially reduce the header dependency as we will be able to include seastar/core/format.hh instead of a more bloated seastar/core/print.hh after bumping up seastar submodule Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21494	2024-11-08 16:43:25 +03:00
Michał Jadwiszczak	f998f027a2	test/boost/secondary_index_test: add test for use after free Reproduces scylladb/scylladb#20772. Add error injection to `raft_server_with_timeouts::read_barrier`, which does 1s sleep before doing the read barrier.	2024-11-08 14:16:19 +01:00
Michał Jadwiszczak	de7b58e8d4	api/raft: use `get_server_with_timeouts().read_barrier()` in coroutines It is unsafe to do `get_server_with_timeouts().read_barrier()` in continuations because `get_server_with_timeouts()` returns raft server by value and it may be deallocated when `read_barrier()` yields, causing use-after-return. Simple workaround is to use the read barrier in coroutine and co_await it. Then the raft server is kept on stack until the read barrier is finished. I've checked all codebase and it looks like the only place where `group0_with_timeouts().read_barrier()` is in continuation, is api/raft.cc. Co-authored-by: Piotr Dulikowski <piodul@scylladb.com>	2024-11-08 14:15:13 +01:00
Botond Dénes	e3e8a94c9a	Merge 'Allow explicitly enabling or disabling tablets when creating a new keyspace' from Benny Halevy Separate the configuration for enabling the tablets feature from the enablement of tablets when creating new keyspaces. This change always enables the TABLETS cluster feature and the tablets logic respectively. The `enable_tablets` config option just controls whether tablets are enabled or disabled by default for new keyspaces. If `enable_tablets` is set to `true`, tablets can be disabled using `CREATE KEYSPACE WITH tablets = { 'enabled': false }` as it is today. If `enable_tablets` is set to `false`, tablets can be enabled using `CREATE KEYSPACE WITH tablets = { 'enabled': true }`. The motivation for this change is to simplify the user experience of using tablets by setting the default for new keyspaces to false amd allowing the user to simply opt-in by using tablets = {enabled: true }. This is not pissible today. The user has to enable tablets by default for all new keyspaces (that use the NetworkTopologyStrategy) and then actively opt-out to use vnodes. * Not required to be backported to OSS versions. May be backported to specific enterprise versions * This PR resubmits https://github.com/scylladb/scylladb/pull/20729 that was reverted in `73b1f66b70` due to https://github.com/scylladb/scylladb/issues/21159 which is now fixed Closes scylladb/scylladb#21451 * github.com:scylladb/scylladb: data_dictionary: keyspace_metadata::describe: print tablets enabled also when defaulted tablets_test: test enable/disable tablets when creating a new keyspace treewide: always allow tablets keyspaces feature_service: prevent enabling both tablets and gossip topology changes alternator: create_keyspace_metadata: enable tablets using feature_service	2024-11-08 09:15:42 +02:00
Michał Chojnowski	35921eb67e	mvcc_test: fix a benign failure of test_apply_to_incomplete_respects_continuity For performance reasons, mutation_partition_v2::maybe_drop(), and by extension also mutation_partition_v2::apply_monotonically(mutation_partition_v2&&) can evict empty row entries, and hence change the continuity of the merged entry. For checking that apply_to_incomplete respects continuity, test_apply_to_incomplete_respects_continuity obtains the continuity of the partition entry before and after apply_to_incomplete by calling e.squashed().get_continuity(). But squashed() uses apply_monotonically(), so in some circumstances the result of squashed() can have smaller continuity than the argument of squashed(), which messes with the thing that the test is trying to check, and causes spurious failures. This patch changes the method of calculating the continuity set, so that it matches the entry exactly, fixing the test failures. Fixes scylladb/scylladb#13757 Closes scylladb/scylladb#21459	2024-11-08 06:08:39 +01:00
Ernest Zaslavsky	029837a4a1	aws_errors: Make error messages more verbose. Add more information to the error messages to make the failure reason clearer. Also add tests to check exceptions propagated from s3 client failure.	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	14f3832749	test: Make the minio proxy randomization re-playable Provide a seed to the proxy randomization, the idea that the `test.py` will initialize the seed from `/dev/urandom` and print the seed when starting, in case some tests failed the dev is supposed to re-play it locally with the same seed (if it didnt repro otherwise) using the `start_s3_proxy.py` and providing it with the aforementioned seed using `--rnd-seed` command line argument	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	0c62635f05	test/boost/s3_test: add error injection scenarios to existing test suite Add variants of existing S3 tests that route through a proxy instead of connecting directly to MinIO. The proxy allows injecting errors to validate error handling and recovery mechanisms under failure conditions.	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	8919e0abab	test: Switch `s3_test` to use proxy Switch `s3_test` to use the S3 proxy which is used to randomly inject retryable S3 errors to test the "retry" part of the S3 client. Fix `put_object` to make it retryable	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	b1e36c868c	test: Add more tests Add tests to cover retries, and retry exhaustion. Also add tests for jumbo upload.	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	7fd1ff8d79	client: Stop returning error on `DELETE` in multipart upload abortion Discard whatever exception is thrown in `abort_upload` when aborting multipart upload since we don't care about cleanly aborting it since there are other means to clean up dangling parts, for example `rclone cleanup` or S3 bucket's Lifecycle Management Policy	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	064a239180	client: Fix sigsegv when retrying Stop moving the `file` into the `make_file_input_stream` since it will try to use it again on retry	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	dc6e4c0d97	client: Add retries Add retries to the s3 client, all retries are coordinated by an instance of `retry_strategy`. In a case of error also parse response body in attempt to retrieve additional and more focused error information as suggested by AWS. See https://docs.aws.amazon.com/AmazonS3/latest/API/ErrorResponses.html. Also move the expected http status check to the `make_s3_error_handler` since the http::client::make_request call is done with `nullopt` - we want to manage all the aws errors handling in s3 client to prevent the http client to validate it and fail before we have a chance to analyze the error properly	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	244635ebd8	client: Adjust `map_s3_client_exception` to return exception instance "Unfuturize" the `map_s3_client_exception` since the retryable client is going to be implemented using coroutines and no `future` is needed here, just to save unnecessary `co_await` on it	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	bd3d4ed417	aws_errors: Change aws_error::parse to return std::optional<> Change aws_error::parse to return std::optional<> to signify that no error was found in the response body	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	58decef509	aws_errors: Add http errors mapping into aws_error Add http errors mapping into aws_error since the retry strategy is going to operate on aws_error and should not be aware of HTTP status codes	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	fa9e8b7ed0	client: Add aws_exception mapping Map aws_exceptions in `map_s3_client_exception`, will be needed in retryable client calls to remap newly added AWS errors to `storage_io_error`	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	54e250a6f1	aws_error: Add `aws_exeption` to carry original `aws_error` Add `aws_exeption` to carry original `aws_error` for proper error handling in retryable s3 client	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	e6ff34046f	aws_errors: Add new error codes Add new S3 related errors, also introduce additional errors to describe pure http errors that has no additional information in the body	2024-11-07 21:01:25 +02:00
Ernest Zaslavsky	8dbe351888	client: Introduce retry strategy Add `retry_strategy` interface and default implementation for exponential back-off retry strategy	2024-11-07 21:01:25 +02:00
Michał Jadwiszczak	7bad8378c7	main,cql_test_env: start group0_service before view_builder In scylladb/scylladb#19745, view_builder was migrated to group0 and since then it is dependent on group0_service. Because of this, group0_service should be initialized/destroyed before/after view_builder. Fixes scylladb/scylladb#20772 Co-authored-by: Dawid Mędrek <dawid.medrek@scylladb.com>	2024-11-07 14:08:11 +01:00
Kamil Braun	c268cf2e33	Merge 'test: rename "cql-pytest" to "cqlpy"' from Nadav Har'El Python and Python developers don't like directory names to include a minus sign, like "cql-pytest". In this patch we rename test/cql-pytest to test/cqlpy, and also change a few references in other code (e.g., code that used test/cql-pytest/run.py) and also references to this test suite in documentation and comments. Arguably, the word "test" was always redundant in test/cql-pytest, and I want to leave the "py" in test/cqlpy to emphasize that it's Python-based tests, contrasting with test/cql which are CQL-request-only approval tests. The second patch in the series fixes a small regression in the test/cqlpy/run script. Fixes #20846 Test organization only, so backports not strictly necessary, but let's do them anyway because otherwise it will make any future backporting of tests in the cqlpy directory more messy than it needs to be. Closes scylladb/scylladb#21446 * github.com:scylladb/scylladb: test/cqlpy: fix "run" script without any parameters test: rename "cql-pytest" to "cqlpy"	2024-11-07 13:26:07 +01:00
Benny Halevy	40928bd886	data_dictionary: keyspace_metadata::describe: print tablets enabled also when defaulted Now that tablets may be explicitly enabled when creating a new keyspace, describe tablets as enabled even when the default initial_tablets==0 is used. Refs https://github.com/scylladb/scylla-enterprise/issues/4860 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-07 13:59:59 +02:00
Benny Halevy	8620d9f672	tablets_test: test enable/disable tablets when creating a new keyspace Test both configuration values for `enable_tablets` and the possibility to explicitly enable or disable tablets, respectively, when creating a keyspace using the `tablets = {'enabled': true\|false}` CREATE KEYSPACE option. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-07 13:57:40 +02:00
Benny Halevy	4b21cca443	treewide: always allow tablets keyspaces With the tablets feature always enabled (Unless gossip toopology changes are forced), the enable_tablets option now controls only the default for newly created keyspaces. Even when set to `false`, tablets are still enabled as a feature and the user may explicitly enable tablets using `CREATE KEYSPACE <name> WITH tablets = {'enabled': true}` Note: best viewed with `git show -w` Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-07 13:57:39 +02:00
Benny Halevy	974b0f2080	feature_service: prevent enabling both tablets and gossip topology changes Tablets require raft consistent topology changes. Therefore, document that they are incompatible in the config help and prevent their usage in `feature_config_from_db_config` Fixes scylladb/scylladb#21075 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-07 13:56:59 +02:00
Benny Halevy	4cf3b683bc	alternator: create_keyspace_metadata: enable tablets using feature_service Rather than using the local configuration option on this node, check the cluster feature instead. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-07 13:56:59 +02:00
Botond Dénes	e21346179c	test/boost/mutation_reader_test: add test for multishard reader buffer hint	2024-11-07 02:47:54 -05:00
Botond Dénes	5c5c77746e	repair/row_level: disable read-ahead The multishard reader's read-ahead was designed to reduce the latency of range scans. But in the case of repair, read-ahead is suspected to contribute significant extra load on the congested streaming semaphore and thus contribute to the subsequent trashing (excessive reader eviction). First off, read-ahead was designed with pages of limited size in mind. Repair can read much more, even for a single repair buffer. This can lead to read-ahead concurrency to continue ramping up, creating and using more and more readers. Secondly, repair is not latency sensitive, so even when working well and there is no congestion, the benefits are negligible. The use of read-ahead is now controllable by the new repair_multishard_reader_enable_read_ahead config item, defaulting to false.	2024-11-07 02:47:54 -05:00
Botond Dénes	a248520201	db/config: introduce repair_multishard_reader_enable_read_ahead Not used yet.	2024-11-07 02:47:54 -05:00
Botond Dénes	36a8756028	readers/multishard: implement the read_ahead flag Don't do read-aheads when read-ahead was not enabled.	2024-11-07 02:47:54 -05:00
Botond Dénes	8938e06ebe	replica/database: make_multishard_streaming_reader(): expose the read_ahead parameter Continuing the previous patch, expose the just added read_ahead parameter of make_multishard_combining>_reader_v2(). Set to read_ahead::yes by all callers, keeping the current default.	2024-11-07 02:47:54 -05:00
Botond Dénes	c6c62deaa5	readers/multishard: add read_ahead parameter And propagate to the reader itself. Not used yet.	2024-11-07 02:47:54 -05:00
Botond Dénes	784f89f585	repair/row_level: set max buffer size on multishard reader The multishard reader is used in the mixed-shard case, when a repair has to read from all other shards. It is very important that cross-shard roundtrips and possible evict-recreate cycles for the shard readers is avoided. For this end, make use of the recently introduced internal buffer hint feature in the multishard reader and set it's buffer size to match that of the row level repair buffer size. The use of the buffer-hint can be controlled with the recently introduced repair_multishard_reader_buffer_hint_size config param.	2024-11-07 02:47:54 -05:00
Botond Dénes	e2344e28b6	replica/database: make_multishard_streaming_reader(): expose buffer_hint parameter Expose the buffer hint functionality added by the previous commits, to callers of make_multishard_streaming_reader(). All callers disable it currently, it will be used in the next patch.	2024-11-07 02:47:46 -05:00
Yaron Kaikov	ef104b7b96	.github/scripts/auto-backport.py: update method to get closed prs `commit.get_pulls()` in PyGithub returns pull requests that are directly associated with the given commit Since in closed PR. the relevant commit is an event type, the backport automation didn't get the PR info for backporting Ref: https://github.com/scylladb/scylladb/issues/18973 Closes scylladb/scylladb#21468	2024-11-07 09:28:46 +02:00
Avi Kivity	9e67649fe5	utils: loading_cache: tighten clock sampling Sample the clock once to avoid the filter returning different results. Range algorithms may use multiple passes, so it's better to return consistent results. Closes scylladb/scylladb#21400	2024-11-07 10:28:01 +03:00
Kefu Chai	50fbab29ca	compaction: remove unused "#include" we don't use `std::list` in compaction/compaction_manager.hh, neither is this header responsible for exposing the declarations in `<list>`. so let's stop `#include` this header. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21436	2024-11-07 10:25:27 +03:00
Avi Kivity	f5489ba4a1	locator: tablet_metadata_guard: forward declare database No need to bring in a heavy databas.hh dependency. Closes scylladb/scylladb#21447	2024-11-07 10:24:35 +03:00
Kefu Chai	ba021f72a6	api: s/mulformatted/malformatted mulformatted was a typo, let's fix it. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21442	2024-11-07 10:07:11 +03:00
Pavel Emelyanov	49949092ad	Merge 'Make s3 client ops use abort source + use in backup task' from Calle Wilund Fixes #20716 Adds optional abort_source to all s3 client operations. If provided, will propagate to actual HTTP client and allow for aborting actual net op. Note: this uses an abort source per call, not a client-local one. This is for two reasons: 1.) The usage pattern of the client object is to create it outside the eventual owning object (task) that hosts the relevant abort source 2.) It is quite possible to want to have different/no abort source for some operation usage. Also adds forward usage of task abort_source in backup tasks upload s3 call, making it more readily abort-able. Closes scylladb/scylladb#21431 * github.com:scylladb/scylladb: backup_task: Use task abort source in s3 client call s3::client: Make operations (individually) abortable	2024-11-07 10:03:25 +03:00
Yaron Kaikov	9d8562caf3	Add conflict_reminder action for backport PR In order not to forget to resolve conflicts in backport PRs, we should add some reminders to the PR author so it will not be forgotten the new action will run twice a week and will send a reminder only for PR opened with conflicts for 3 days or more Fixes: https://github.com/scylladb/scylladb/issues/21448 Closes scylladb/scylladb#21449	2024-11-07 06:55:37 +02:00
Calle Wilund	0db4b9fd94	backup_task: Use task abort source in s3 client call Fixes #20716 Propagates abort source in task object to actual network call, thus making the upload workload more quickly abortable. v2: Fix test to handle two versions after each other	2024-11-06 15:20:23 +00:00
Nadav Har'El	1fd7b797c7	test/cqlpy: fix "run" script without any parameters A recent improvement to test/cqlpy/run to add the "--release" option broke the ability to run this script it without any options (no test name, etc.). This patch fixes this case. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-11-06 16:48:36 +02:00
Nadav Har'El	8c215141a1	test: rename "cql-pytest" to "cqlpy" Python and Python developers don't like directory names to include a minus sign, like "cql-pytest". In this patch we rename test/cql-pytest to test/cqlpy, and also change a few references in other code (e.g., code that used test/cql-pytest/run.py) and also references to this test suite in documentation and comments. Arguably, the word "test" was always redundant in test/cql-pytest, and I want to leave the "py" in test/cqlpy to emphasize that it's Python-based tests, contrasting with test/cql which are CQL-request-only approval tests. Fixes #20846 Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-11-06 16:48:36 +02:00
Kefu Chai	6efde20939	utils/to_string: do not include fmt/ostream.h to_string.hh does not use this header, neither is it obliged to expose the content of this header. so, let's remove this include. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21440	2024-11-06 17:21:29 +03:00
Botond Dénes	3c25e6fcb4	db/config: introduce enable_repair_multishard_reader_buffer_hint Allows enabling/disabling the multishard reader buffer hint optimization. Not wired yet.	2024-11-06 08:51:00 -05:00
Botond Dénes	b052c5df62	readers/multishard: multishard_reader: pass hint to shard_reader Calculate a buffer fill hint and pass it to shard_reader_v2::fill_buffer(), so the underlying buffer-fill can be optimized to avoid multiple cross shard round-trips, as well as possible evict-recreate cycles. The buffer hint mechanism is opt-in, enabled via the new multishard_reader_buffer_hint parameter.	2024-11-06 08:51:00 -05:00
Botond Dénes	912b4dfba3	readers/multishard: shard_reader_v2::fill_reader_buffer(): respect the hint When the hint is provided, respect it: make sure the returned buffer is of the requested size, stopping early if the stop_token is seen. To reduce the amount of possible eviction-recreate cycles while the buffer is filled, disable auto-pause for the duration of the fill_reader_buffer() call. For this purpose, auto_pause_disable_guard is added to evictable_reader_v2.	2024-11-06 08:51:00 -05:00
Botond Dénes	8d5283f036	readers/multishard: propagate fill_buffer_hint to shard_reader:fill_reader_buffer() The hint will tell the shard reader exactly how much data to produce, to avoid multiple cross-shard round-trips and possible evict-recreate cycles. The hint is neither used yet or calculated yet, this is coming in the next patches.	2024-11-06 08:51:00 -05:00
Botond Dénes	ee7ecb9155	readers/multishard: shard_reader: extract buffer-fill into its own method It is about to get a bit more complicated, so worth to extract into a method so it can be shared by the two call-sites.	2024-11-06 08:51:00 -05:00
Tomasz Grabiec	f7d35d535e	Merge 'bytes_ostream: replace boost ranges with std ranges' from Avi Kivity To reduce the dependency load, replace boost ranges with std::ranges. Cleanup; no backport. Closes scylladb/scylladb#21450 * github.com:scylladb/scylladb: bytes_ostream: replace boost ranges with std ranges bytes_ostream: extract fragment_iterator into namespace scope	2024-11-06 14:01:27 +01:00
Yaron Kaikov	77604b4ac7	.github/script/auto-backport.py: push backport PR to `scylladbbot` fork Since Scylla is a public repo, when we create a fork, it doesn't fork the team and permissions (unlike private repos where it does). When we have a backport PR with conflicts, the developers need to be able to update the branch to fix the conflicts. To do so, we modified the logic of the backport automation as follows: - Every backport PR (with and without conflicts) will be open directly on the `scylladbbot` fork repo - When there are conflicts, an email will be sent to the original PR author with an invitation to become a contributor in the `scylladbbot` fork with `push` permissions. This will happen only once if Auther is not a contributor. - Together with sending the invite, all backport labels will be removed and a comment will be added to the original PR with instructions - The PR author must add the backport labels after the invitation is accepted Fixes: https://github.com/scylladb/scylladb/issues/18973 Closes scylladb/scylladb#21401	2024-11-06 14:29:37 +02:00
David Garcia	a072478f4f	docs: enable tooltips Updates the theme to the latest version to enable tooltips and modifies the db_options.tmpl to show the new role in action. Closes scylladb/scylladb#21324	2024-11-06 14:09:28 +02:00
Andrei Chekun	afd1fc8e9f	test.py: Add pytest-xdist to the toolchain Add new dependency pytest-xdist to the toolchain. This will allow executing boost and unit tests from pytest in parallel, reducing the time needed for the run. Closes scylladb/scylladb#21222	2024-11-06 14:09:01 +02:00
Botond Dénes	0ad32c153d	Merge 'test_tablets: add rack decommission test cases' from Benny Halevy test_tablets: add rack decommission test cases Test scenarios where decommissioing a compelte rack should succeed, and reproduce scylladb/scylladb#19475 where decommissioning a rack would fail since the number of remaining racks is insufficient to satisfy the replication factor, even though the number of nodes is sufficient, enshrining this behavior. Refs scylladb/scylladb#19475 * This PR adds unit tests and improves an error message. No backport required. Closes scylladb/scylladb#20747 * github.com:scylladb/scylladb: tablet_allocator: improve error message when unable to find replicas when draining test_tablets: add rack decommission test cases topology_experimental_raft/test_tablets: get_tablet_count_per_shard_for_host: move shards_count param to be last test/pylib: ServerInfo: add datacenter and rack attributes test: everywhere: drop unused imports of ServerInfo	2024-11-06 14:07:47 +02:00
Calle Wilund	3321820c67	s3::client: Make operations (individually) abortable Refs #20716 Adds optional abort_source to all s3 client operations. If provided, will propagate to actual HTTP client and allow for aborting actual net op. Note: this uses an abort source per call, not a client-local one. This is for two reasons: 1.) The usage pattern of the client object is to create it outside the eventual owning object (task) that hosts the relevant abort source 2.) It is quite possible to want to have different/no abort source for some operation usage.	2024-11-05 14:23:24 +00:00
Avi Kivity	8fb6d98ba3	Merge "various gossiper code cleanups" from Gleb * 'gleb/gossip-cleanup-v3' of github.com:scylladb/scylla-dev: gossiper: start failure_detector_loop on shard 0 only gossiper: use 1 seconds instead of 1000 milliseconds gossiper: remove unused code gossiper: co-routinize do_send_ack2_msg gossiper: do not needlessly call get_endpoint_state_ptr in handle_major_state_change gossiper: fix weird logic in get_live_members gossiper: drop unneeded this-> gossiper: fold get_or_create_endpoint_state into my_endpoint_state gossiper: co-routinize do_send_ack_msg	2024-11-05 15:31:58 +02:00
Avi Kivity	baaa92c6f5	bytes_ostream: replace boost ranges with std ranges Have fragment_iterator support iterator_concept for compatibility with std ranges, and switch from boost iterator_range to std::ranges::subrange.	2024-11-05 14:50:38 +02:00
Avi Kivity	cb026c347e	bytes_ostream: extract fragment_iterator into namespace scope C++ concept evaluation rules clash with nested class definition rules with the result that evaluating concepts about the nested class within the enclosing class doesn't work. Extract bytes_ostream::fragment_iterator to avoid that.	2024-11-05 14:43:49 +02:00
Avi Kivity	4dab2473a2	Merge 'treewide: trade boost's any_of and all_of for std's any_of and all_of' from Kefu Chai now that we are allowed to use C++23. we now have the luxury of using `std::ranges::all_of` and `std::ranges::any_of` in this change, we replace `boost::algorithm::all_of` and `boost::algorithm::any_of` with `std::ranges::all_of` and `std::ranges::any_of` respectively. to reduce the dependency to boost for better maintainability, and leverage standard library features for better long-term support. this change is part of our ongoing effort to modernize our codebase and reduce external dependencies where possible. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#21411 * github.com:scylladb/scylladb: treewide: s/boost::algorithm::any_of/std::ranges::any_of/ treewide: s/boost::algorithm::all_of/std::ranges::all_of/	2024-11-05 12:48:24 +02:00
Piotr Dulikowski	7f17894c88	Merge 'cql3: Allow for describing CDC log tables' from Dawid Mędrek In the past, DESC SCHEMA would produce create statements for both the base and the log table. That was incorrect as the log table is automatically created alongside the base one. That was solved in scylladb/scylladb@9ab57b1 (scylladb/scylladb#18467). The mentioned changes implemented the following solution: * DESC SCHEMA/KEYSPACE/TABLE would still print a create statement for the CDC base table, * DESC SCHEMA/KEYSPACE would start printing an alter statement for the CDC log table. That statement would ensure that the restored log table has the same parameters as the original one, * DESC TABLE <base table> would behave as DESC SCHEMA/KEYSPACE, i.e. it would print a create statement for the base table and an alter statement for the log table, * DESC TABLE <log table> would result in an error. While that solution was good and behaved correctly in the context of restoring the schema, it had one flaw: describe statement aren't only used as a means for producing a backup; they also serve an informative purpose to learn about the schema, e.g. to learn what parameters a specific table uses. Because we didn't allow for describing CDC log tables, the user couldn't look them up directly via a describe statement -- they had to describe the base table for that. Attempting to describe a log table ended with an error, e.g.: ``` $ DESC TABLE ks.t_scylla_cdc_log; ks.t_scylla_cdc_log is a cdc log table and it cannot be described directly. Try `DESC TABLE ks.t` to describe cdc base table and it's log table. ``` In these changes, we allow for describing CDC log tables again. The semantics of the first three bullets above remains unchanged, but we impose new behavior for DESC TABLE <log table>: * When the user executes DESC TABLE <log table>, a create statement will be returned, treating the table as if it were a regular one, * The create statement will be wrapped in CQL comment markers. The rationale for the second bullet is that although we want to give the user a means to look into the structure and options of a CDC log table, the returned statement is not supposed to be ever executed by them. We want to minimize the risk of that. An example of the behavior after the change: ``` $ DESC TABLE ks.t_scylla_cdc_log; /* Do NOT execute this statement! It's only for informational purposes. A CDC log table is created automatically when the base is created. CREATE TABLE ks.t_scylla_cdc_log ( "cdc$stream_id" blob, "cdc$time" timeuuid, "cdc$batch_seq_no" int, "cdc$end_of_batch" boolean, "cdc$operation" tinyint, "cdc$ttl" bigint, p int, PRIMARY KEY ("cdc$stream_id", "cdc$time", "cdc$batch_seq_no") ) WITH CLUSTERING ORDER BY ("cdc$time" ASC, "cdc$batch_seq_no" ASC) AND bloom_filter_fp_chance = 0.01 AND caching = {'enabled': 'false', 'keys': 'NONE', 'rows_per_partition': 'NONE'} AND comment = 'CDC log for ks.t' AND compaction = {'class': 'TimeWindowCompactionStrategy', 'compaction_window_size': '60', 'compaction_window_unit': 'MINUTES', 'expired_sstable_check_frequency_seconds': '1800'} AND compression = {'sstable_compression': 'org.apache.cassandra.io.compress.LZ4Compressor'} AND crc_check_chance = 1 AND default_time_to_live = 0 AND gc_grace_seconds = 0 AND max_index_interval = 2048 AND memtable_flush_period_in_ms = 0 AND min_index_interval = 128 AND speculative_retry = '99.0PERCENTILE'; / ``` We also extend the developer documentation regarding DESCRIBE statements on CDC tables. Fixes scylladb/scylladb#21235 Backport: these changes are an enhancement, so not needed. Closes scylladb/scylladb#21228 github.com:scylladb/scylladb: docs/dev: Document semantics of describing CDC tables cql3: Allow for describing CDC log tables	2024-11-05 10:06:13 +01:00
Pavel Emelyanov	440c1e3e3f	error_injection: Remove unused inject(sleep, then invoke) overload The overload was introduced by `a8b14b0227` (utils: add timeout error injection with lambda), but is only used by the test nowadays. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21377	2024-11-05 09:56:08 +02:00
Yaniv Michael Kaul	4c5e102aee	node_exporter: use fewer collectors Remove unused / less useful collectors by default. While it doesn't seem to reduce memory usage, it may reduce potential performance or security issues in the future. This is what we are left with (snippet of log when loading node exporter manually with the changed command line): ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:111 level=info msg="Enabled collectors" ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=arp ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=bonding ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=conntrack ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=cpu ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=cpufreq ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=diskstats ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=dmi ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=edac ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=entropy ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=filefd ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=filesystem ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=interrupts ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=loadavg ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=mdadm ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=meminfo ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=netclass ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=netdev ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=netstat ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=nvme ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=os ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=pressure ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=schedstat ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=selinux ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=sockstat ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=softnet ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=stat ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=textfile ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=time ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=timex ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=uname ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=vmstat ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=watchdog ts=2024-11-03T15:41:06.855Z caller=node_exporter.go:118 level=info collector=xfs Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Improvement, no need to backport. Closes scylladb/scylladb#21419	2024-11-05 10:41:09 +03:00
Avi Kivity	b292aeecac	replica: query.hh: drop dependency on database.hh database.hh has large fan-in and therefore can trigger a lot of recompilations if included. Replace with smaller dependencies. Closes scylladb/scylladb#21424	2024-11-05 10:40:33 +03:00
Pavel Emelyanov	a98b57212e	Merge 'lang, .github: remove unused includes, add more directories to CLEANER_DIR' from Kefu Chai in this series: - remove unused `#include` in "lang" subdirectory - add index and lang to CLEANER_DIR --- cleanup and improvements in the CI, hence no need to backport. Closes scylladb/scylladb#21437 * github.com:scylladb/scylladb: .github: add index and lang to CLEANER_DIR lang: remove unused "#includes"	2024-11-05 10:37:04 +03:00
Kefu Chai	f1d4812ad6	test: lib: rest_client: use isinstance() over type() in addition to the inheritance support, `isinstance()` is also the recommended way to check for types by PEP8. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21438	2024-11-05 10:36:31 +03:00
Kefu Chai	59eb2ab119	treewide: s/boost::algorithm::any_of/std::ranges::any_of/ now that we are allowed to use C++23. we now have the luxury of using `std::ranges::any_of`. in this change, we replace `boost::algorithm::any_of` with `std::ranges::any_of` to reduce the dependency to boost for better maintainability, and leverage standard library features for better long-term support. this change is part of our ongoing effort to modernize our codebase and reduce external dependencies where possible. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-11-05 14:06:09 +08:00
Kefu Chai	f8bb1c64f1	treewide: s/boost::algorithm::all_of/std::ranges::all_of/ now that we are allowed to use C++23. we now have the luxury of using `std::ranges::all_of`. in this change, we replace `boost::algorithm::all_of` with `std::ranges::all_of` to reduce the dependency to boost for better maintainability, and leverage standard library features for better long-term support. this change is part of our ongoing effort to modernize our codebase and reduce external dependencies where possible. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-11-05 14:05:24 +08:00
Kefu Chai	e651b6dc69	.github: add index and lang to CLEANER_DIR also explain why we don't run the cleaner against the "idl" subdirectory. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-11-05 10:01:04 +08:00
Kefu Chai	ee2a9419b3	lang: remove unused "#includes" these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-11-05 10:01:04 +08:00
Avi Kivity	ee92784098	serialization: replace boost::type with std::type_identity Recently, seastar rpc started accepting std::type_identity in addition to boost::type as a type marker (while labeling the latter with an ominous deprecation warning). Reduce our depedendency on boost by switching to std::type_identity.	2024-11-05 00:43:27 +01:00
Avi Kivity	075b13597d	serializer: drop dependency on boost ranges The call to boost::range::for_each is easily replaced with ranged for. Closes scylladb/scylladb#21422	2024-11-04 17:48:17 +02:00
Gleb Natapov	2dbae78542	gossiper: start failure_detector_loop on shard 0 only failure_detector_loop does nothing on all other shards.	2024-11-04 17:15:06 +02:00
Gleb Natapov	323b04137d	gossiper: use 1 seconds instead of 1000 milliseconds	2024-11-04 17:15:06 +02:00
Gleb Natapov	0cb4c71846	gossiper: remove unused code	2024-11-04 17:15:06 +02:00
Gleb Natapov	0e4f149dee	gossiper: co-routinize do_send_ack2_msg	2024-11-04 17:15:06 +02:00
Gleb Natapov	1fbac54fb8	gossiper: do not needlessly call get_endpoint_state_ptr in handle_major_state_change The code calls for get_endpoint_state_ptr several times instead of using the result of the first call. Change it.	2024-11-04 17:15:06 +02:00
Gleb Natapov	e2cf93abb9	gossiper: fix weird logic in get_live_members The code adds a node to a set and then removes it if a condition is met. Add to the set if the condition is not met instead. Note that the original set never has local endpoint (it is only added locally), so the code is equivalent.	2024-11-04 17:14:55 +02:00
Benny Halevy	caedcf20c6	tablet_allocator: improve error message when unable to find replicas when draining Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-04 14:54:54 +02:00
Benny Halevy	d8be1cafb5	test_tablets: add rack decommission test cases Test scenarios where decommissioing a compelte rack should succeed, and reproduce scylladb/scylladb#19475 where decommissioning a rack would fail since the number of remaining racks is insufficient to satisfy the replication factor, even though the number of nodes is sufficient, enshrining this behavior. Refs scylladb/scylladb#19475 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-04 14:54:13 +02:00
Avi Kivity	b706e3e9e4	Merge 'sstables/index_reader: avoid unnecessary index page reads in single-partition reads' from Michał Chojnowski Terminology note: in the context of this series, "index page" means an contiguous segment of the index file starting (inclusive) at a key corresponding to a summary entry and ending (exclusive) before the key corresponding to the next summary entry. "Index pages" are not related to filesystem pages. --- In a single-partition read, if the searched partition key is the first key in its index page, we start scanning the index for that key starting at the previous index page (inclusive), even though we could start directly from the key's page. Similarly, if the searched partition key is absent from the sstable and lies after all other keys in its appropriate page, we additionally scan the next page, even though it's known from the summary that it can't possibly contain the key. Those cases are wasteful. It's worse than it might seem at first glance. When partitions are small, only a small fraction of search keys fulfills those conditions (i.e. "first key in its page" or "an absent key greater than the last key in its page"), so the waste doesn't matter much. But when partitions are big enough, every index page contains only one partition key (and a promoted index for that partition), which directly means that all search keys fulfill the conditions, which means that total index reading work is two times bigger than what it should be. In addition, there is a secondary performance bug which, when the aforementioned conditions are fulfilled, causes additional I/O to happen past the index reads which are actually parsed and used. In effect, the index I/O in single-partition reads might be not just doubled, but even tripled (that's for IOPS — throughput might be multiplied even more), all because of a slight inaccuracy in the edge cases. This series fixes those inefficiencies by tightening the edge cases and ensuring that single-partition reads always read only a single index page. Here's an example where we query the first row (i.e. `LIMIT 1`) of a certain partition key, in a table with large (1 MB) promoted indexes. Before the patch, the lookup of the lower bound involves 3 serialized disk reads (as described above) to subsequent index pages, and even the lookup of the upper bound involves 2 disk reads: ``` Execute CQL3 query Parsing a statement [shard 0] Processing a statement for authenticated user: anonymous [shard 0] Executing read query (reversed false) [shard 0] Creating read executor for token -1297921881139976049 with all: [127.11.11.1] targets: [127.11.11.1] repair decision: NONE [shard 0] Creating never_speculating_read_executor - speculative retry is disabled or there are no extra replicas to speculate with [shard 0] read_data: querying locally [shard 0] Start querying singular range {{-1297921881139976049, pk{00023130}}} [shard 0] [reader concurrency semaphore user] admitted immediately [shard 0] [reader concurrency semaphore user] executing read [shard 0] Reading key {-1297921881139976049, pk{00023130}} from sstable ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Data.db [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 38359040 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 38391808 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 38359040, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 38391808, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 39370752 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 39403520 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 39370752, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 39403520, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 40378368 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 40411136 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 40378368, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 40411136, successfully read 32768 bytes [shard 0] upper_bound_cache_only({position: clustered, ckp{}, 1}): no upper bound [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 40378368 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 40411136 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 40378368, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 40411136, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 41390080 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 41422848 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 41390080, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 41422848, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Data.db: scheduling bulk DMA read of size 21926 at offset 819200 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Data.db: finished bulk DMA read of size 21926 at offset 819200, successfully read 24576 bytes [shard 0] Page stats: 1 partition(s), 0 static row(s) (0 live, 0 dead), 1 clustering row(s) (1 live, 0 dead), 0 range tombstone(s) and 0 cell(s) (0 live, 0 dead) [shard 0] Querying is done [shard 0] Done processing - preparing a result [shard 0] Request complete ``` After the patch, the lookup of each bound involves 1 read: ``` Execute CQL3 query Parsing a statement [shard 0] Processing a statement for authenticated user: anonymous [shard 0] Executing read query (reversed false) [shard 0] Creating read executor for token -1297921881139976049 with all: [127.11.11.1] targets: [127.11.11.1] repair decision: NONE [shard 0] Creating never_speculating_read_executor - speculative retry is disabled or there are no extra replicas to speculate with [shard 0] read_data: querying locally [shard 0] Start querying singular range {{-1297921881139976049, pk{00023130}}} [shard 0] [reader concurrency semaphore user] admitted immediately [shard 0] [reader concurrency semaphore user] executing read [shard 0] Reading key {-1297921881139976049, pk{00023130}} from sstable ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Data.db [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 39370752 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 39403520 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 39370752, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 39403520, successfully read 32768 bytes [shard 0] upper_bound_cache_only({position: clustered, ckp{}, 1}): no upper bound [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 40378368 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: scheduling bulk DMA read of size 32768 at offset 40411136 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 40378368, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Index.db: finished bulk DMA read of size 32768 at offset 40411136, successfully read 32768 bytes [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Data.db: scheduling bulk DMA read of size 21926 at offset 819200 [shard 0] ./workdir_01/data/ks/t-536c31f09a9c11efbd5082a6aa3e8d0c/me-3gky_0v18_3rgjk2dsjae431s4uz-big-Data.db: finished bulk DMA read of size 21926 at offset 819200, successfully read 24576 bytes [shard 0] Page stats: 1 partition(s), 0 static row(s) (0 live, 0 dead), 1 clustering row(s) (1 live, 0 dead), 0 range tombstone(s) and 0 cell(s) (0 live, 0 dead) [shard 0] Querying is done [shard 0] Done processing - preparing a result [shard 0] Request complete ``` Doesn't have to be backported, since the problem only affects performance, not correctness, and it has been present since forever. Closes scylladb/scylladb#20897 * github.com:scylladb/scylladb: index_reader: remove a piece of misguided code involved in single-partition reads index_reader: in single-partition reads, don't read more than one page index_reader: fix unnecessary reads of preceding index pages	2024-11-04 14:28:27 +02:00
Avi Kivity	2531dc2d80	schema_registry: stop including replica/database.hh database.hh is a hotspot that changes often (or its dependencies do). Avoid including it to reduce recompilations. Closes scylladb/scylladb#21407	2024-11-04 13:16:27 +01:00
Benny Halevy	9ff614da9f	topology_experimental_raft/test_tablets: get_tablet_count_per_shard_for_host: move shards_count param to be last Prepare for the next comit that will add a version accepting a list of servers: `get_tablet_count_per_shard_for_hosts` for which we want `shards_per_node` to be last and have a default value. Also, fix the type hint for `full_tables`, as it had a syntax error, using `:` instead of `,`. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-04 14:11:30 +02:00
Benny Halevy	0c1e85b6e3	test/pylib: ServerInfo: add datacenter and rack attributes Set to "DEFAULT_DC" and "DEFAULT_RACK" by default. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-04 14:11:30 +02:00
Benny Halevy	efa64cb92a	test: everywhere: drop unused imports of ServerInfo Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-04 14:11:30 +02:00
Avi Kivity	7cb1ad8c87	Merge 'compaction_manager: stop_tasks, stop_ongoing_compactions: ignore errors' from Benny Halevy stop() methods, like destructors must always succeed, and returning errors from them is futile as there is nothing else we can do with them by continue with shutdown. stop_ongoing_compactions, in particular, currently returns the status of stopped compaction tasks from `stop_tasks`, but still all tasks must be stopped after it, even if they failed, so assert that and ignore the errors. Fixes scylladb/scylladb#21159 * Needs backport to 6.2 and 6.1, as commit `8cc99973eb` causes handles storage that might cause compaction tasks to fail and eventually terminate on shudown when the exceptions are thrown in noexcept context in the deferred stop destructor body Closes scylladb/scylladb#21299 * github.com:scylladb/scylladb: compaction_manager: stop: await _stop_future if engaged compaction_manager: really_do_stop: assert that no tasks are left behind compaction_manager: stop_tasks, stop_ongoing_compactions: ignore errors compaction/compaction_manager: stop_tasks(): unlink stopped tasks compaction/compaction_manager: make _tasks an intrusive list	2024-11-04 13:54:16 +02:00
Gleb Natapov	3b7d9fddbc	gossiper: drop unneeded this->	2024-11-04 12:02:51 +02:00
Gleb Natapov	501b8f6984	gossiper: fold get_or_create_endpoint_state into my_endpoint_state my_endpoint_state() is the only called of get_or_create_endpoint_state() and calling it is the only thing the function does anyway.	2024-11-04 12:02:51 +02:00
Gleb Natapov	300cbcebf6	gossiper: co-routinize do_send_ack_msg	2024-11-04 12:02:51 +02:00
Pavel Emelyanov	f3f956841f	sstables: Remove unused mp_row_consumer_m::range_tombstone_start It's only used by its operator<< so remove it as well Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21380	2024-11-03 16:40:02 +02:00
Avi Kivity	704ea9d3b4	Merge 'api: Remove foreach_column_family() helper' from Pavel Emelyanov There's a whole lot of helpers and wrappers in api/ that help handlers manipulate keyspaces and tables. One of those is foreach_column_family which calls the provided callable on a table on each shard. There's exactly the same (but a bit more flexible) helper nearby. While at it, this helper gets a better name. Closes scylladb/scylladb#21398 * github.com:scylladb/scylladb: api: Rename set_tables -> for_tables_on_all_shards api: Remove foreach_column_family() helper	2024-11-03 15:46:27 +02:00
Avi Kivity	856489ded1	cql3: remove unused request_validations methods These methods are not used and therefore removed. Closes scylladb/scylladb#21392	2024-11-03 13:17:32 +02:00
Benny Halevy	6cce67bec8	compaction_manager: stop: await _stop_future if engaged The current condition that consults the compaction manager state for awaiting `_stop_future` works since _stop_future is assigned after the state is set to `stopped`, but it is incidental. What matters is that `_stop_future` is engaged. While at it, exchange _stop_future with a ready future so that stop() can be safely called multiple times. And dropped the superfluous co_return. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-03 10:53:35 +02:00
Benny Halevy	a7a55298ea	compaction_manager: really_do_stop: assert that no tasks are left behind stop_ongoing_compactions now ignores any errors returned by tasks, and it should leave no task left behind. Assert that here, before the compaction_manager is destroyed. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-03 10:53:34 +02:00
Benny Halevy	c08ba8af68	compaction_manager: stop_tasks, stop_ongoing_compactions: ignore errors stop() methods, like destructors must always succeed, and returning errors from them is futile as there is nothing else we can do with them but continue with shutdown. Leaked errors on the stop path may cause termination on shutdown, when called in a deferred action destructor. Fixes scylladb/scylladb#21298 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-03 10:52:58 +02:00
Botond Dénes	d8500472b3	compaction/compaction_manager: stop_tasks(): unlink stopped tasks Stopped tasks currently linger in _tasks until the fiber that created the task is scheduled again and unlinks the task. This window between stop and remove prevents reliable checks for empty _tasks list after all tasks are stopped. Unlink the task early so really_do_stop() can safely check for an empty _tasks list (next patch).	2024-11-03 10:17:11 +02:00
Botond Dénes	e942c074f2	compaction/compaction_manager: make _tasks an intrusive list _tasks is currently std::list<shared_ptr<compaction_task_executor>>, but it has no role in keeping the instances alive, this is done by the fibers which create the task (and pin a shared ptr instance). This lends itself to an intrusive list, avoiding that extra allocation upon push_back(). Using an intrusive list also makes it simpler and much cheaper (O(1) vs. O(N)) to remove tasks from the _tasks list. This will be made use of in the next patch. Code using _task has to be updated because the value_type changes from shared_ptr<compaction_task_executor> to compaction_task_executor&.	2024-11-03 10:17:11 +02:00
Avi Kivity	39b55bd3a0	Update seastar submodule * seastar f821bda19...fba36a3d1 (13): > build: do not include -DBoost_TEST_DYN_LINK in seastar_testing_cflags > doc: compatibility: update the notes on supported GCC versions > docker: bump up to clang {18,19} and gcc {13,14} > rpc: optimize small tuple deserialization > rpc: switch rpc::type from boost to std > thread: do not use fortify source > build: suppress CMake warning about CMP0057 > core/units: remove space before literal identifier > signal.md: describe auto signal handling > build: persist Seastar options in SeastarConfig.cmake > sharded.hh: seperate invoke_on decls from defs > test: Add perf test for http client > gate: check: mark as const Closes scylladb/scylladb#21390	2024-11-02 13:58:45 +02:00
Botond Dénes	19a43b5859	Merge 'repair: Reduce hints and batchlog flush' from Asias He The hints and batchlog flush requests are issued to all nodes for each repair request when tombstone_gc repair mode is used. The amount of such flush requests is high when all nodes in the cluster run repair. It is observed it takes a long time, up to 15s, for a repair request to finish such a flush request. To reduce overhead of the flush, each node caches the flush and only executes the real flush when some time has passed. It is safe to do so before the real flush_time is returned. Repair uses the smallest flush_time from peers as the repair time. The nice thing about the cache on the receiver side is that all senders can hit the cache. It is better than cache on the sender side. A slightly smaller flush_time compared to the real flush time will be used with the benefits of significantly dropped hints and batchlog flush. The tradeoff is reasonable. Fixes #20259 Performance improvement. No backports. Closes scylladb/scylladb#20260 * github.com:scylladb/scylladb: test/test_repair.py: Add test_batchlog_flush_in_repair repair: Reduce hints and batchlog flush db/batchlog_manager: Add add_delay_to_batch_replay db/batchlog_manager: Add get_last_replay db/batchlog_manager: wire in batchlog_replay_cleanup_after_replays db/config: introduce batchlog_replay_cleanup_after_replays db/batchlog_manager: do_batch_log_replay(): add cleanup flag	2024-11-01 14:23:27 +02:00
Pavel Emelyanov	292fd52a60	Merge 'utils: chunked_vector: various constructor improvements' from Avi Kivity Optimize the various constructors a little, and add an std::from_range_t constructor. Minor improvement, so no backports. Closes scylladb/scylladb#21399 * github.com:scylladb/scylladb: utils: chunked_vector: add from_range_t constructor utils: chunked_vector: optimize initializer_list constructor utils: chunked_vector: iterator constructor: copy spanwise utils: chunked_vector: reserve for forward iterators, not just random access iterators, on construction	2024-11-01 15:02:56 +03:00
Botond Dénes	4bafaee523	Merge 'tasks: improve task_manager::lookup_virtual_task' from Aleksandra Martyniuk Currently, to find the operation with given id, all operations tracked by a virtual task are listed. This isn't necessary, since we only need info regarding one particular operation. Add a method to check whether a virtual task tracks the operation with the given id. No backport needed Closes scylladb/scylladb#20769 * github.com:scylladb/scylladb: tasks: delete virtual_task::get_ids method as it is unused tasks: improve task_manager::lookup_virtual_task	2024-11-01 13:44:04 +02:00
Kefu Chai	1b8446f92d	compaction: fix the indent in `38ce2c605d`, we left a TODO for reindent the code. in this change, we reindent the code to address this TODO. Refs `38ce2c605d` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21383	2024-11-01 12:55:47 +03:00
Avi Kivity	b5e46077df	sstables: generation_type: replace boost ranges with std ranges Reduce dependency load. Closes scylladb/scylladb#21402	2024-11-01 12:45:24 +03:00
Pavel Emelyanov	d6169630a4	api: Rename set_tables -> for_tables_on_all_shards The former name is not extremely descriptive, hopefully the latter one is better in this sense. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-11-01 12:15:01 +03:00
Pavel Emelyanov	822758dffd	api: Remove foreach_column_family() helper There's a whole lot of helpers and wrappers in api/ that help handlers manipulate keyspaces and tables. One of those is foreach_column_family which calls the provided callable on a table on each shard. There's exactly the same (but a bit more flexible) set_table() helper nearby. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-11-01 12:13:35 +03:00
Botond Dénes	0ee0dd3ef4	Merge 'Collect and report backup progress' from Pavel Emelyanov Task manager GET /status method returns two counters that reflect task progress -- total and completed. To make caller reason about their meaning, additionally there's progress_units field next to those counters. This patch implements this progress report for backup task. The units are bytes, the total counter is total size of files that are being uploaded, and the completed counter is total amount of bytes successfully sent with PUT requests. To get the counters, the client::upload_file() is extended to calculate those. fixes #20653 Closes scylladb/scylladb#21144 * github.com:scylladb/scylladb: backup_task: Report uploading progress s3/client: Account upload progress for real s3/client: Introduce upload_progress s3: Extract client_fwd.hh	2024-11-01 10:57:12 +02:00
Kefu Chai	64122b3df3	treewide: s/boost::transform/std::ranges::transform/ now that we are allowed to use C++23. we now have the luxury of using `std::ranges::transform`. in this change, we: - replace `boost::transform` with `std::ranges::transform` - update affected code to work with `std::ranges::transform` to reduce the dependency to boost for better maintainability, and leverage standard library features for better long-term support. this change is part of our ongoing effort to modernize our codebase and reduce external dependencies where possible. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21318	2024-11-01 08:15:14 +02:00
Avi Kivity	8c67f9b42e	cql3: util: remove unneeded boost/range includes from header files The includes are redistributed to the source files that need them. Closes scylladb/scylladb#21391	2024-10-31 23:49:44 +01:00
Nadav Har'El	ee2d75b088	Merge 'Generalize "breakpoint" type of error injection' from Pavel Emelyanov This pattern is -- if requested (by test) suspend code execution until requestor (the test) explicitly wakes it up. For that the injected place should inject a lambda that is called with so called "handler" at hand and try to read message from the handler. In many cases the inner lambda additionally prints a message into logs that tests waits upon to make sure injection was stepped on. In the end of the day this "breakpoint" is injected like ``` co_await inject("foo", [] (auto& handler) { log.info("foo waiting"); co_await handler.wait_for_message(timeout); }); ``` This PR makes breakpoints shorter and more unified, like this ``` co_await inject("foo", wait_for_message(timeout)); ``` where `wait_for_message` is a wrapper structure used to pick new `inject()` overload. Closes scylladb/scylladb#21342 * github.com:scylladb/scylladb: sstables: Use inject(wait_for_message_overload) treewide,error_injection: Use inject(wait_for_message) and fix tests treewide,error_injection: Use inject(wait_for_message) overload error_injection: Add inject() overload with wait_for_message wrapper	2024-10-31 21:56:27 +02:00
Avi Kivity	6a9852d47b	utils: chunked_vector: add from_range_t constructor std::ranges::to<> has a little protocol with containers. Implement it to get optimized construction. Similar to the iterator pair constructor, if the range's size can be obtained (even with an O(N) algorithm), favor that to avoid reallocations. Copy elements spanwise to promote optimization to memcpy when possible.	2024-10-31 19:32:16 +02:00
Avi Kivity	b2769403d2	utils: chunked_vector: optimize initializer_list constructor Delegate to the previously optimized iterator-pair constructor.	2024-10-31 18:10:14 +02:00
Avi Kivity	0a81be4321	utils: chunked_vector: iterator constructor: copy spanwise Instead of copying element-by-element, copy contiguous spans. This is much faster if the input is a span and the constructor is trivial, since the whole thing translates to a memcpy. Make the two branches constexpr to reduce work for the compiler in optimizing the other branch away.	2024-10-31 18:10:08 +02:00
Avi Kivity	4653430c8e	utils: chunked_vector: reserve for forward iterators, not just random access iterators, on construction For a forward iterator, prefer a two pass algorithm to first count the number of elements, reserver, then copy the elements, to a single pass algorithm that involves reallocation and copying.	2024-10-31 17:55:42 +02:00
Kefu Chai	673b107ffa	github: use GithubException when appropriate `Exception` could be too general, what we really care about is `GithubException`. so let's catch the latter instead for better readability. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21364	2024-10-31 18:21:29 +03:00
Kefu Chai	f8221b960f	test: route S3 mock server messages through logger The S3 mock server (introduced in `5a96549c`) currently prints its status messages directly to stdout, which can be distracting when reviewing test results. For example: ```console $ ./test.py --verbose --mode debug object_store/test_backup::test_simple_backup Found 1 tests. Starting S3 mock server on ('127.226.51.1', 2012) ================================================================================ [N/TOTAL] SUITE MODE RESULT TEST ------------------------------------------------------------------------------ [1/1] object_store debug [ PASS ] object_store.test_backup.1 5.99s Stopping S3 mock server ------------------------- CPU utilization: 6.5% ``` Move these messages to use proper logging to give developers more control over their visibility: - Make logger parameter mandatory in MockS3Server constructor - Route "Stopping S3 mock server" message through the provided logger - Add --log-level option to the standalone mock server launcher The message is now hidden: ```console $ ./test.py --verbose --mode debug --save-log-on-success object_store/test_backup::test_simple_backup Found 1 tests. ================================================================================ [N/TOTAL] SUITE MODE RESULT TEST ------------------------------------------------------------------------------ [1/1] object_store debug [ PASS ] object_store.test_backup.1 6.25s ------------------------------------------------------------------------------ CPU utilization: 5.5% ``` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21384	2024-10-31 18:21:29 +03:00
Benny Halevy	78ceaeabca	compaction_manager: compaction_disabled: return true if not in compaction_state When a compaction_group is removed via `compaction_manager::remove`, it is erase from `_compaction_state`, and therefore compaction is definitely not enabled on it. This triggers an internal error if tablets are cleaned up during drop/truncate, which checks that compaction is disabled in all compaction groups. Note that the callers of `compaction_disabled` aren't really interested in compaction being actively disabled on the compaction_group, but rather if it's enabled or not. A follow-up patch can be consider to reverse the logic and expose `compaction_enabled` rather than `compaction_disabled`. Fixes scylladb/scylladb#20060 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#21378	2024-10-31 18:21:29 +03:00
Dawid Mędrek	495c1188e9	docs/dev: Document semantics of describing CDC tables	2024-10-31 11:25:19 +01:00
Dawid Mędrek	39e0513e1b	cql3: Allow for describing CDC log tables In the past, DESC SCHEMA would produce create statements for both the base and the log table. That was incorrect as the log table is automatically created alongside the base one. That was solved in scylladb/scylladb@9ab57b1 (scylladb/scylladb#18467). The mentioned changes implemented the following solution: * DESC SCHEMA/KEYSPACE/TABLE would still print a create statement for the CDC base table, * DESC SCHEMA/KEYSPACE would start printing an alter statement for the CDC log table. That statement would ensure that the restored log table has the same parameters as the original one, * DESC TABLE <base table> would behave as DESC SCHEMA/KEYSPACE, i.e. it would print a create statement for the base table and an alter statement for the log table, * DESC TABLE <log table> would result in an error. While that solution was good and behaved correctly in the context of restoring the schema, it had one flaw: describe statement aren't only used as a means for producing a backup; they also serve an informative purpose to learn about the schema, e.g. to learn what parameters a specific table uses. Because we didn't allow for describing CDC log tables, the user couldn't look them up directly via a describe statement -- they had to describe the base table for that. Attempting to describe a log table ended with an error, e.g.: ``` $ DESC TABLE ks.t_scylla_cdc_log; ks.t_scylla_cdc_log is a cdc log table and it cannot be described directly. Try `DESC TABLE ks.t` to describe cdc base table and it's log table. ``` In these changes, we allow for describing CDC log tables again. The semantics of the first three bullets above remains unchanged, but we impose new behavior for DESC TABLE <log table>: * When the user executes DESC TABLE <log table>, a create statement will be returned, treating the table as if it were a regular one, * The create statement will be wrapped in CQL comment markers. The rationale for the second bullet is that although we want to give the user a means to look into the structure and options of a CDC log table, the returned statement is not supposed to be ever executed by them. We want to minimize the risk of that. An example of the behavior after the change: ``` $ DESC TABLE ks.t_scylla_cdc_log; /* Do NOT execute this statement! It's only for informational purposes. A CDC log table is created automatically when the base is created. CREATE TABLE ks.t_scylla_cdc_log ( "cdc$stream_id" blob, "cdc$time" timeuuid, "cdc$batch_seq_no" int, "cdc$end_of_batch" boolean, "cdc$operation" tinyint, "cdc$ttl" bigint, p int, PRIMARY KEY ("cdc$stream_id", "cdc$time", "cdc$batch_seq_no") ) WITH CLUSTERING ORDER BY ("cdc$time" ASC, "cdc$batch_seq_no" ASC) AND bloom_filter_fp_chance = 0.01 AND caching = {'enabled': 'false', 'keys': 'NONE', 'rows_per_partition': 'NONE'} AND comment = 'CDC log for ks.t' AND compaction = {'class': 'TimeWindowCompactionStrategy', 'compaction_window_size': '60', 'compaction_window_unit': 'MINUTES', 'expired_sstable_check_frequency_seconds': '1800'} AND compression = {'sstable_compression': 'org.apache.cassandra.io.compress.LZ4Compressor'} AND crc_check_chance = 1 AND default_time_to_live = 0 AND gc_grace_seconds = 0 AND max_index_interval = 2048 AND memtable_flush_period_in_ms = 0 AND min_index_interval = 128 AND speculative_retry = '99.0PERCENTILE'; */ ``` Fixes scylladb/scylladb#21235	2024-10-31 11:25:19 +01:00
Wojciech Mitros	88ab8db944	mv: run view building in streaming scheduling group View building is an expensive process that takes a long time to complete. During the build, it's impact on other work should be minimized, even at the expense of slightly slowing it down. Instead, view building is currently performed in the the same scheduling group (gossip) as other high-priority tasks, in particular raft processing, which slows it down, making races more likely and increasing the number of retries that need to be done. While view building is still initiated in the gossip group (as it's the result of adding a view, which is a schema change), in this patch the bulk of the view building work is moved to a low-priority, maintenance scheduling group (named "streaming" after its main use case). Additionally, a test is added, where we make sure that the scheduling group is the one most used when building a view. Fixes https://github.com/scylladb/scylladb/issues/21232 Closes scylladb/scylladb#21326	2024-10-31 10:13:20 +01:00
Nadav Har'El	7572c483b1	test/topology_experimental_raft: fix flaky test Today, each test function in test/topology_experimental_raft creates a cluster in the beginning of the test and drops it at the end of the function. This is very inefficient if you hope (like I do) to write many small and pinpointed test functions instead of large test functions that test 20 unrelated things. Trying to propose a way to change this sad state of affairs, in test_alternator.py I created a fixture "alternator3" which I hoped could be used in multiple tests that need a 3-node Alternator cluster. Currently only one test uses this fixture. Unfortunately, it turns out the alternator3 fixture is broken, and led to flaky test runs (sometimes the test using alternator3 picked up an existing cluster instead of starting with an empty cluster, and failed). These problems cannot be completely fixed at the current state of the framework. The framework does not currently allow keeping a 3-node cluster between test functions, while also allowing other test functions to create different clusters. The specific flakiness we saw could be fixed by adding a missing before_test() call, but in the future we would need to ensure that all the test functions that use it are contiguous in the test file, and I don't see how we can (or want to) ensure this. So at this point I am giving up and withdrawing this proposal until the developers of the topology test framework make this one of their design goals. Since there was only one test using this fixture, removing it should make no performance or correctness difference - it should just fix the flakiness. Fixes scylladb/scylladb#21322. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#21370	2024-10-31 10:12:26 +01:00
Calle Wilund	c4361037f7	cql_test_env/gossip: Prevent double shutdown call crash Fixes scylladb/scylladb#21159 When an exception is thrown in sstable write etc such that storage_manager::isolate is initiated, we start a shutdown chain for message service, gossip etc. These are synced (properly) in storage_manager::stop, but if we somehow call gossiper::shutdown outside the normal service::stop cycle, we can end up running the method simultaneously, intertwined (missing the guard because of the state change between check and set). We then end up co_awaiting an invalid future (_failure_detector_loop_done) - a second wait. Fixed by a.) Remove superfluous gossiper::shutdown in cql_test_env. This was added in `20496ed`, ages ago. However, it should not be needed nowadays. b.) Ensure _failure_detector_loop_done is always waitable. Just to be sure. Closes scylladb/scylladb#21379	2024-10-31 10:11:20 +01:00
Nadav Har'El	d3f09638f0	Merge 'compound_compat: replace use of boost ranges with std ranges' from Avi Kivity Replace use of boost::ranges::join() with another construct, as it has no std replacement, and replace other uses with their std equivalent, in order to reduce dependency load. Code cleanup - no backport. Closes scylladb/scylladb#21382 * github.com:scylladb/scylladb: compound_compat: replace use of boost ranges with std ranges compound_compat: simplify seriakization of ka/la sstables static cell names	2024-10-31 10:16:41 +02:00
Nadav Har'El	65e29f28bd	Merge 'gms: remove unused #includes ' from Kefu Chai these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#21374 * github.com:scylladb/scylladb: .github: add gms to iwyu's CLEANER_DIR gms: remove unused `#include`s	2024-10-31 09:06:37 +02:00
Kefu Chai	2498e37a2f	mutation_writer,streaming: use reader_consumer_v2 type when appropriate The `reader_consumer_v2` type (`std::function<future<> (mutation_reader)>`) is defined alongside `mutation_reader` in `mutation_reader.hh`. before this change, we sometimes use `std::function<future<> (mutation_reader)>` directly when defining a consumer parameter or a consumer variable. in this change, we improve maintainability by: - Reducing duplicate function type declarations - Centralizing the consumer type definition - Making future signature updates easier to implement Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21369	2024-10-31 07:17:47 +02:00
Avi Kivity	907da210b6	compound_compat: replace use of boost ranges with std ranges To reduce the dependency load, replace use of boost ranges with the std equivalent. Files that lost the indirect boost dependency have it added as a direct dependency.	2024-10-30 19:58:07 +02:00
Avi Kivity	982cebc1f6	compound_compat: simplify seriakization of ka/la sstables static cell names compound_compat is used for serializing ka/la sstables static cell names. Since we can no longer write such sstabkes, the function is used only in some tests. Reduce the use of boost::range::join(): it has no direct equivalent in std (std::views::concat is in C++26), and it is slow due to the need to type-erase. Instead of using boost::range::join, extend the vector used to hold the empty clustering key a bit more, and copy the view representing the static cell name into into it.	2024-10-30 19:19:57 +02:00
Kefu Chai	d3a6931b14	.github: add gms to iwyu's CLEANER_DIR to avoid future violations of include-what-you-use. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-30 23:01:34 +08:00
Kefu Chai	52ec315ffd	gms: remove unused `#include`s these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-30 23:01:34 +08:00
Pavel Emelyanov	c16369323b	sstables: Use inject(wait_for_message_overload) This place could be in the pre-previous patch, it just can use the overload, but it seemengly has a bug. It prints _two_ messages -- that the injection handler was suspended and that it was woken up. The bug is in the 2nd message -- it's printed without waiting for the message, so it likely gets printed before wakeup itself. It seems that no tests care about it though. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-30 16:53:33 +03:00
Pavel Emelyanov	39cb93be3c	treewide,error_injection: Use inject(wait_for_message) and fix tests This is continuation of previous patch, this time also update tests that wait for specific message in logs (to make sure injection handler was called and paused the code execution). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-30 16:53:33 +03:00
Pavel Emelyanov	7d8cc3ccc2	treewide,error_injection: Use inject(wait_for_message) overload Many places want to inject a handler that waits for external kick. Now there's convenience inject() method overload for this. It will result in extra messages in logs, but so far no code/test cares about it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-30 16:53:33 +03:00
Pavel Emelyanov	c1432f3657	error_injection: Add inject() overload with wait_for_message wrapper The wrapper object denotes that injection should run a handler and wait_for_message() on it. Wrapper carries the timeout used to call the mentioned method. It's currently unused, next patches will start enjoing it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-30 16:53:33 +03:00
Dawid Mędrek	b984488552	cql3: Rename `SALTED HASH` to `HASHED PASSWORD` Cassandra 4.1 announced a new option to create a role with: `HASHED PASSWORD`. Example: ``` CREATE ROLE bob WITH HASHED PASSWORD = 'hashed_password'; ``` We've already introduced another option following the same semantics: `SALTED HASH`; example: ``` CREATE ROLE bob WITH SALTED HASH = 'salted_hash'; ``` The change hasn't made it to any release yet, so in this commit we rename it to `HASHED PASSWORD` to be compatible with Cassandra. Additionally, we adjust existing tests to work against Cassandra too. Fixes scylladb/scylladb#21350 Closes scylladb/scylladb#21352	2024-10-30 14:07:58 +02:00
Aleksandra Martyniuk	bc5b1f9a5d	tasks: delete virtual_task::get_ids method as it is unused	2024-10-30 12:25:47 +01:00
Aleksandra Martyniuk	9b5d69ae96	tasks: improve task_manager::lookup_virtual_task Currently, lookup_virtual_task gets the list of ids of all operations tracked by a virtual task and checks whether it contains given id. The list of all ids isn't required and the check whether one particular operation id is tracked by the virtual task may be quicker than listing all operations. Add virtual_task::contains method and use it in lookup_virtual_task.	2024-10-30 12:24:38 +01:00
Kefu Chai	d81ed5adb4	compaction: explain make_interpose_consumer() in compaction strategy Add documentation to clarify the purpose and behavior of make_interpose_consumer() in the compaction_strategy_impl class. This method is crucial for building layered processing pipelines but its semantics were previously undocumented. The added documentation explains how: - It decorates end consumers with additional processing steps - It enables construction of processing pipelines - The original consumer's semantics are preserved This improves code maintainability by making the pipeline construction pattern more apparent to developers. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21336	2024-10-30 13:22:00 +03:00
Tomasz Grabiec	f3869dadc6	Merge 'compound: replace boost ranges with std ranges' from Avi Kivity Continue standardization on std::ranges. Since compound contains a custom iterator, we first have to upgrade it to C++20 iterator concepts. Cleanup / minor refactoring, so no backport. Closes scylladb/scylladb#21320 * github.com:scylladb/scylladb: compound: replace boost ranges with std ranges compound: upgrade iterator to be an std::forward_iterator	2024-10-30 11:02:51 +01:00
Asias He	73806f66a5	test/test_repair.py: Add test_batchlog_flush_in_repair It checks batchlog flush request cache in repair.	2024-10-30 11:10:39 +08:00
Asias He	b3b3e880d3	repair: Reduce hints and batchlog flush The hints and batchlog flush requests are issued to all nodes for each repair request when tombstone_gc repair mode is used. The amount of such flush requests is high when all nodes in the cluster run repair. It is observed it takes a long time, up to 15s, for a repair request to finish such a flush request. To reduce overhead of the flush, each node caches the flush and only executes the real flush when the cahce time has passed. It is safe to do so because the real flush_time is returned. Repair uses the smallest flush_time returned from peers as the repair time. The nice thing about the cache on the receiver side is that all senders can hit the cache. It is better than cache on the sender side. A slightly smaller flush_time compared to the real flush time will be used with the benefits of significantly dropped hints and batchlog flush. The trade-off looks reasonable. Tests: 2 nodes, with 1s batchlog delay: Before: Repair nr_repairs=20 cache_time_in_ms=0 total_repair_duration=40.04245328903198 After: Repair nr_repairs=20 cache_time_in_ms=5000 total_repair_duration=1.252073049545288 Fixes #20259	2024-10-30 11:07:57 +08:00
Asias He	f8ad78ba1e	db/batchlog_manager: Add add_delay_to_batch_replay It is used to simulate slow replay.	2024-10-30 11:07:57 +08:00
Asias He	fed9b54664	db/batchlog_manager: Add get_last_replay It is used to get the time when the last replay is executed.	2024-10-30 11:07:57 +08:00
Botond Dénes	3361542e84	db/batchlog_manager: wire in batchlog_replay_cleanup_after_replays After the specified amount of replays, trigger a cleanup: flush batchlog table memtables. This allows the cleanup to happen on a configurable interval, instead of on every batchlog replay attempt, which might be too much.	2024-10-30 11:07:57 +08:00
Botond Dénes	1635525526	db/config: introduce batchlog_replay_cleanup_after_replays Not used yet.	2024-10-30 11:07:57 +08:00
Botond Dénes	169c74346d	db/batchlog_manager: do_batch_log_replay(): add cleanup flag Add a flag controlling whether cleanup (memtable flush) will be done after the replay. This is to allow repair to opt out from cleanup -- when many concurrenty repairs are running, there can be storms of calles to do_batch_log_replay(), which will be mostly no-op, but they will all attempt to flush the memtable to clean-up after themselves. This is unnecessary and introduces latency to repairs, best to leave the cleanup to the periodic batch-log replay.	2024-10-30 11:07:57 +08:00
Avi Kivity	73b1f66b70	Revert "Merge 'Allow explicitly enabling or disabling tablets when creating a new keyspace' from Benny Halevy" This reverts commit `c286434e4c`, reversing changes made to `6712fcc316`. The commit causes memtable_test to be very flaky in debug mode. Specifically, subtests test_exceptions_in_flush_on_sstable_open and test_exceptions_in_flush_on_sstable_write).	2024-10-30 00:55:29 +02:00
Avi Kivity	b9df3aec12	gdb: avoid @classmethod/@property combinations The @classmethod/@property combination was deprecated in Python 3.11 and removed[1] in Python 3.13. It's used in scylla-gdb.py, breaking it with Python 3.13. To fix, just make all users (size_t and _vptr_type) top-level functions. The definitions are all identical and don't need to be in class scope. [1] https://docs.python.org/3.13/library/functions.html#classmethod Closes scylladb/scylladb#21349	2024-10-29 19:37:07 +02:00
Gleb Natapov	cc7f25062a	topology coordinator: take a copy of a replication state in raft_topology_cmd_handler Current code takes a reference and holds it past preemption points. And while the state itself is not suppose to change the reference may become stale because the state is re-created on each raft topology command. Fix it by taking a copy instead. This is a slow path anyway. Fixes: scylladb/scylladb#21220 Closes scylladb/scylladb#21316	2024-10-29 15:47:43 +01:00
Avi Kivity	020ccbd76a	Merge 'utils: cached_file: Mark permit as awaiting on page miss' from Tomasz Grabiec Otherwise, the read will be considered as on-cpu during promoted index search, which will severely underutlize the disk because by default on-cpu concurrency is 1. I verified this patch on the worst case scenario, where the workload reads missing rows from a large partition. So partition index is cached (no IO) and there is no data file IO (relies on https://github.com/scylladb/scylladb/pull/20522). But there is IO during promoted index search (via cached_file). Before the patch this workload was doing 4k req/s, after the patch it does 30k req/s. The problem is much less pronounced if there is data file or partition index IO involved because that IO will signal read concurrency semaphore to invite more concurrency. Fixes #21325 Closes scylladb/scylladb#21323 * github.com:scylladb/scylladb: utils: cached_file: Mark permit as awaiting on page miss utils: cached_file: Push resource_unit management down to cached_file	2024-10-29 16:15:21 +02:00
Kamil Braun	36cc3bcc90	test: test_crash_coordinator_before_streaming: enable TRACE for `raft_topology` logger Issue scylladb/scylladb#21114 reported that sometimes during the test we timeout when waiting for node to restart after it was killed. Preliminary investigation showed that the node appears to be hanging inside `topology_state_load`, while holding `token_metadata` lock, which prevents `join_topology` from progressing. Enable TRACE level logging for `raft_topology` so we get more accurate info where inside `topology_state_load` the hang happens, once the problem reproduces again in CI. Closes scylladb/scylladb#21247	2024-10-29 12:46:47 +02:00
Kefu Chai	54d438168a	build: cmake: explicitly mark convenience libraries as STATIC before this change, these [convenience libraries](https://www.gnu.org/software/automake/manual/html_node/Libtool-Convenience-Libraries.html) were implicitly built as static libraries by default, but weren't explicitly marked as STATIC in CMake. While this worked with default settings, it could cause issues if `BUILD_SHARED_LIBS` is enabled. So before we are ready for building these components as shared libraries, let's mark all convenience libraries as STATIC for consistency and to prevent potential issues before we properly support shared library builds. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21274	2024-10-29 10:22:19 +01:00
Yaron Kaikov	94a9efbf1c	github: add script for backports automation instead of Mergify Adding an auto-backport.py script to handle backport automation instead of Mergify. The rules of backport are as follows: * Merged or Closed PRs with any backport/x.y label (one or more) and promoted-to-master label * Backport PR will be automatically assigned to the original PR author * In case of conflicts the backport PR will be open in the original autoor fork in draft mode. This will give the PR owner the option to resolve conflicts and push those changes to the PR branch (Today in Scylla when we have conflicts, the developers are forced to open another PR and manually close the backport PR opened by Mergify) * Fixing cherry-pick the wrong commit SHA. With the new script, we always take the SHA from the stable branch * Support backport for enterprise releases (from Enterprise branch) Fixes: https://github.com/scylladb/scylladb/issues/18973 Closes scylladb/scylladb#21302	2024-10-29 10:04:30 +02:00
Pavel Emelyanov	25ae3d0aed	backup_task: Report uploading progress Do it by passing reference to s3::upload_progress_monitor object that sits on task impl itself. Different files' uploads would then update the monitor with their sizes and uploaded counters. The structure is reported by get_progress() method. Unit size is set to be bytes. Test is updated. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-29 08:40:35 +03:00
Pavel Emelyanov	2efcfc13e8	s3/client: Account upload progress for real Before upload starts file size is checked, so this is the place that updates progress.total counter. Uploading a file happens by reading unit_size bytes from file input stream and writing the buffer into http body writer stream. This is the place to update progress.uploaded counter. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-29 08:38:39 +03:00
Pavel Emelyanov	51e03b1025	s3/client: Introduce upload_progress This is a structure with "total" and "uploaded" counters that's passed by user to client::upload_file() method so that client would update it with the progress. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-29 08:38:39 +03:00
Pavel Emelyanov	f9a5e02b53	s3: Extract client_fwd.hh This is to export some simple structures to users without the need to include client.hh itself (rather large already) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-29 08:38:39 +03:00
Avi Kivity	49d3e281d6	Merge 'Sanitize /system/highest_supported_sstable_version API endpoint' from Pavel Emelyanov Its handler dereferences long chain of objects to get to the value it needs. There's shorter way. Also, the endpoint in question is not unregistered on stop. Closes scylladb/scylladb#21279 * github.com:scylladb/scylladb: api: Make get_highest_supported_sstable_version use proper service api: Move system::get_highest_supported_sstable_version set/unset api: Scaffold for sstables-format-selector	2024-10-28 21:42:41 +02:00
Pavel Emelyanov	b09bb6bc19	error_injection: Re-use enter() code in inject() overloads Most of inject() overloads check if the injection is enabled, then optionally clear the one-shot one, then do the injection. Everything but doing the injection is implemented in the enter() method, it's perfectly worth re-using one. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21285	2024-10-28 21:37:20 +02:00
Kefu Chai	7610b907c6	build: include subdirectory rules in compilation database merge Previously in `e65185ba`, when merging Seastar's and ScyllaDB's compilation databases, the "prefix" parameter in merge-compdb.py was too restrictive. It only included build rules for files with "CMakeFiles" prefix, excluding source files in subdirectories like `apps/iotune/CMakeFiles/app_iotune.dir/iotune.cc.o`. In this change, we change the prefix parameter to an empty string to include all source files whose object files are located under build directories, regardless of their path structure. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21312	2024-10-28 21:34:21 +02:00
Avi Kivity	c286434e4c	Merge 'Allow explicitly enabling or disabling tablets when creating a new keyspace' from Benny Halevy Separate the configuration for enabling the tablets feature from the enablement of tablets when creating new keyspaces. This change always enables the TABLETS cluster feature and the tablets logic respectively. The `enable_tablets` config option just controls whether tablets are enabled or disabled by default for new keyspaces. If `enable_tablets` is set to `true`, tablets can be disabled using `CREATE KEYSPACE WITH tablets = { 'enabled': false }` as it is today. If `enable_tablets` is set to `false`, tablets can be enabled using `CREATE KEYSPACE WITH tablets = { 'enabled': true }`. The motivation for this change is to simplify the user experience of using tablets by setting the default for new keyspaces to false amd allowing the user to simply opt-in by using tablets = {enabled: true }. This is not pissible today. The user has to enable tablets by default for all new keyspaces (that use the NetworkTopologyStrategy) and then actively opt-out to use vnodes. * Not required to be backported to OSS versions. May be backported to specific enterprise versions Closes scylladb/scylladb#20729 * github.com:scylladb/scylladb: data_dictionary: keyspace_metadata::describe: print tablets enabled also when defaulted tablets_test: test enable/disable tablets when creating a new keyspace treewide: always allow tablets keyspaces feature_service: prevent enabling both tablets and gossip topology changes alternator: create_keyspace_metadata: enable tablets using feature_service	2024-10-28 21:33:17 +02:00
Nadav Har'El	6712fcc316	test/cql-pytest: add option to run cql-pytes tests against specific release This patch adds the option "--release <version>" to test/cql-pytest/run, which downloads the pre-compiled Scylla release with the given version number and runs the tests against that version. For example, it can be used to demonstrate that #15559 was indeed a regression between 2022.1 and 2022.2, by running a recently-added test against these two old versions: test/cql-pytest/run --release 2022.1 --runxfail \ test_prepare.py::test_duplicate_named_bind_marker_prepared test/cql-pytest/run --release 2022.2 --runxfail \ test_prepare.py::test_duplicate_named_bind_marker_prepared The first run passes, the second fails - showing the regression. The Scylla releases are downloaded from ScyllaDB's S3 bucket (downloads.scylladb.com). They are saved in the build/ directory (e.g., build/2022.2.9), and if that directory is not removed, when "run --release" requests the same version again, the previous download is reused. Release numbers can look like: * 5.4.7 * 5.4 (will get the latest in the 5.4 branch, e.g., 5.4.7) * 5.4.0~rc2 (a prerelease) * 2021.1.9 (Enterprise release) * 2023.1 (latest in this branch, Enterprise release) Fixes #13189 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#19228	2024-10-28 21:29:44 +02:00
Kefu Chai	f3dee5b636	build: enable CMAKE_CXX_EXTENSIONS explicitly before this change, Seastar enables CXX_EXTENSIONS in its own build rules. but it does not expose it to the parent project. but scylladb's CMake building system respect seastar's .pc file and includes the cflags exposed by it. without this change, scylladb included "-std=c++23" from seastar, and "-std=gnu++23" from itself. this is both confusing and inconsistent with the build rules generated by `configure.py`. in this change, we explicitly set `CMAKE_CXX_EXTENSIONS` when creating Seastar's building rules, so that it can populate this setting to its .pc file. in this way, we don't have two different options for specifying the C++ standard when building scylladb with CMake. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21311	2024-10-28 21:23:04 +02:00
Kefu Chai	8b80ef3290	build: Remove GCC ARM warning workaround (originally added in `193d1942`) The workaround was initially added to silence warnings on GCC < 6.4 for ARM platforms due to a compiler bug (gcc.gnu.org/bugzilla/show_bug.cgi?id=77728). Since our codebase now requires modern GCC versions for coroutine support, and the bug was fixed in GCC 6.4+, this workaround is no longer needed. Refs `193d1942f2` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21308	2024-10-28 21:19:56 +02:00
Avi Kivity	94c21e5c05	Merge 'sstables: Reduce amount of I/O for clustering-key-bounded reads from large partitions' from Tomasz Grabiec Single-row reads from large partition issue 64 KiB reads to the data file, which is equal to the default span of the promoted index block in the data file. If users would want to increase selectivity of the index to speed up single-row reads, this won't be effective. The reason is that the reader uses promoted index to look up the start position in the data file of the read, but end position will in practice extend to the next partition, and amount of I/O will be determined by the underlying file input stream implementation and its read-ahead heuristics. By default, that results in at least 2 IOs 32KB each. There is already infrastructure to lookup end position based on upper bound of the read, in anticipation for sharing the promoted index cache, but it's not effective becasue it's a non-populating lookup and the upper bound cursor has its own private cached_promoted_index, which is cold when positions are computed. It's non-populating on purpose, to avoid extra index file IO to read upper bound. In case upper bound is far-enough from the lower bound, this will only increase the cost of the read. The solution employed here is to warm up the lower bound cursor's cache before positions are computed, and use that cursor for non-populating lookup of the upper bound. We use the lower bound cursor and the slice's lower bound so that we read the same blocks as later lower-bound slicing would, so that we don't incur extra IO for cases where looking up upper bound is not worth it, that is when upper bound is far from the lower bound. If upper bound is near lower bound, then warming up using lower bound will populate cached_promoted_index with blocks which will allow us to locate the upper bound block accurately. This is especially important for single-row reads, where the bounds are around the same key. In this case we want to read the data file range which belongs to a single promoted index block. It doesn't matter that the upper bound is not exactly the same. They both will likely lie in the same block, and if not, binary search will bring adjacent blocks into cache. Even if upper bound is not near, the binary search will populate the cache with blocks which can be used to narrow down the data file range somewhat. Fixes #10030. The change was tested with perf-fast-forward. I populated the data set with `column_index_size_in_kb` set to 1 scylla perf-fast-forward --populate --run-tests=large-partition-slicing --column-index-size-in-kb=1 Test run: build/release/scylla perf-fast-forward --run-tests=large-partition-select-few-rows -c1 --keep-cache-across-test-cases --test-case-duration=0 This test issues two reads of subsequent keys from the middle of a large partition (1M rows in total). The first read will miss in the index file page cache, the second read will hit. Notice that before the change, the second read issued 2 aio requests worth of 64KiB in total. After the change, the second read issued 1 aio worth of 2 KiB. That's because promoted index block is larger than 1 KiB. I verified using logging that the data file range matches a single promoted index block. Also, the first read which misses in cache is still faster after the change. Before: ``` running: large-partition-select-few-rows on dataset large-part-ds1 Testing selecting few rows from a large partition: stride rows time (s) iterations frags frag/s mad f/s max f/s min f/s avg aio aio (KiB) blocked dropped idx hit idx miss idx blk c hit c miss c blk allocs tasks insns/f cpu 500000 1 0.009802 1 1 102 0 102 102 21.0 21 196 2 1 0 1 1 0 0 0 568 269 4716050 53.4% 500001 1 0.000321 1 1 3113 0 3113 3113 2.0 2 64 1 0 1 0 0 0 0 0 116 26 555110 45.0% ``` After: ``` running: large-partition-select-few-rows on dataset large-part-ds1 Testing selecting few rows from a large partition: stride rows time (s) iterations frags frag/s mad f/s max f/s min f/s avg aio aio (KiB) blocked dropped idx hit idx miss idx blk c hit c miss c blk allocs tasks insns/f cpu 500000 1 0.009609 1 1 104 0 104 104 20.0 20 137 2 1 0 1 1 0 0 0 561 268 4633407 43.1% 500001 1 0.000217 1 1 4602 0 4602 4602 1.0 1 2 1 0 1 0 0 0 0 0 110 26 313882 64.1% ``` Backports: none, not a regression Closes scylladb/scylladb#20522 * github.com:scylladb/scylladb: perf: perf_fast_forward: Add test case for querying missing rows perf-fast-forward: Allow overriding promoted index block size perf-fast-forward: Test subsequent key reads from the middle in test_large_partition_select_few_rows perf-fast-forward: Allow adding key offset in test_large_partition_select_few_rows perf-fast-forward: Use single-partition reads in test_large_partition_select_few_rows sstables: bsearch_clustered_cursor: Add more tracing points sstables: reader: Log data file range sstables: bsearch_clustered_cursor: Unify skip_info logging sstables: bsearch_clustered_cursor: Narrow down range using "end" position of the block sstables: bsearch_clustered_cursor: Skip even to the first block test: sstables: sstable_3_x_test: Improve failure message sstables: mx: writer: Never include partition_end marker in promoted index block width sstables: Reduce amount of I/O for clustering-key-bounded reads from large partitions sstables: clustered_cursor: Track current block	2024-10-28 21:13:23 +02:00
Tomasz Grabiec	0f2101b055	utils: cached_file: Mark permit as awaiting on page miss Otherwise, the read will be considered as on-cpu during promoted index search, which will severely underutlize the disk because by default on-cpu concurrency is 1. I verified this patch on the worst case scenario, where the workload reads missing rows from a large partition. So partition index is cached (no IO) and there is no data file IO. But there is IO during promoted index search (via cached_file). Before the patch this workload was doing 4k req/s, after the patch it does 30k req/s. The problem is much less pronounced if there is data file or index file IO involved because that IO will signal read concurrency semaphore to invite more concurrency.	2024-10-28 19:54:58 +01:00
Tomasz Grabiec	868f5b59c4	utils: cached_file: Push resource_unit management down to cached_file It saves us permit operations on the hot path when we hit in cache. Also, it will lay the ground for marking the permit as awaiting later.	2024-10-28 19:49:58 +01:00
Avi Kivity	d3dae09316	compound: replace boost ranges with std ranges Standardize on the standard range library. The serialize_value(initializer_list) overload is disambiguated not to call itself. Apparently it wasn't called before. Since std::ranges::subrange does not provide operator==, replace it with std::ranges::equals().	2024-10-28 18:35:41 +02:00
Avi Kivity	61d7f1f6a5	compound: upgrade iterator to be an std::forward_iterator compound::iterator isn't far from a forward_iterator, and if we want to use it with std::ranges, we have to upgrade it. This is because std::ranges::subrange() only provides front() for forward ranges, and we do use this front(). Boost apparently isn't as strict. To make it a forward_range, we have to drop operator-> and make operator* return a value (similar to std::views::tranform), since forward iterators require that pointers and references be stable, and this iterator returns a pointer to one of its members. We also add an iterator_concept member to declare the compatibility to std::ranges.	2024-10-28 17:16:36 +02:00
Kamil Braun	101c1d50f0	Merge 'fix nodetool status to show zero-token nodes' from Abhinav Kumar Jha In the current scenario, the nodetool status doesn’t display information regarding zero token nodes. For example, if 5 nodes are spun by the administrator, out of which, 2 nodes are zero token nodes, then nodetool status only shows information regarding the 3 non-zero token nodes. This commit intends to fix this issue by leveraging the “/storage_service/host_id ” API and adding appropriate logic in scylla-nodetool.cc to support zero token nodes. A test is also added in nodetool/test_status.py to verify this logic. This test fails without this commit’s zero token node support logic, hence verifying the behavior. This PR fixes a bug. Hence we need to backport it. Backporting needs to be done only to 6.2 version, since earlier versions don't support zero token nodes. Fixes: scylladb/scylladb#19849 Fixes: scylladb/scylladb#17857 Closes scylladb/scylladb#20909 * github.com:scylladb/scylladb: fix nodetool status to show zero-token nodes test: move `wait_for_first_completed` to pylib/util.py token_metadata: rename endpoint_to_host_id_map getter and add support for joining nodes	2024-10-28 12:19:36 +01:00
Kefu Chai	9f8adcd207	backup_task: track the first failure uploading sstables before this change, we only record the exception returned by `upload_file()`, and rethrow the exception. but the exception thrown by `update_file()` not populated to its caller. instead, the exceptional future is ignored on pupose -- we need to perform the uploads in parallel. this is why the task is not marked fail even if some of the uploads performed by it fail. in this change, we - coroutinize `backup_task_impl::do_backup()`. strictly speaking, this is not necessary to populate the exception. but, in order to ensure that the possible exception is captured before the gate is closed, and to reduce the intentation, the teardown steps are performed explicitly. - in addition to note down the exception in the logging message, we also store it in a local variable, which it rethrown before this function returns. Fixes scylladb/scylladb#21248 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21254	2024-10-28 12:54:27 +03:00
Tzach Livyatan	1878af9399	Update os-support-info.rst - add CentOS ScyllaDB support RHEL 9 and derivatives, including CentOS 9. Fix https://github.com/scylladb/scylladb/issues/21309 Closes scylladb/scylladb#21310	2024-10-28 10:02:31 +02:00
Kefu Chai	8ac471b74b	dht: do not include unused headers in `8d1b3223`, we removed some unused "#include"s, but we failed to address all of them in "dht" subdirectory. and the unaddressed "#include"s are identified by the iwyu workflow. in this change, we address the leftovers. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21291	2024-10-28 09:58:42 +02:00
Anna Stuchlik	44a807f5bc	doc: improve the README file in the docs folder This commit improves the README file so that it's more helpful to documentation contributors. Especially, it: - Adds the link to the prerequisites. - Add information on troubleshooting (checking the links, headings, etc.) - Removes the section on creating a knowledge base article, as we no longer promote adding KBs in favor of creating a coherent documentation set. Fixes https://github.com/scylladb/scylladb/issues/21257 Closes scylladb/scylladb#21262	2024-10-28 09:55:40 +02:00
Anna Stuchlik	212eb204a7	doc: set 6.2 as the latest stable version This commit updates the configuration for ScyllaDB documentation so that: - 6.2 is the latest version. - 6.2 is removed from the list of unstable versions. It must be merged when ScyllaDB 6.2 is released. In addition, this commit uncomments the redirections that should be applied when version 6.2 is the latest stable version (which will happen when this commit is merged). No backport is required. Closes scylladb/scylladb#21133	2024-10-28 09:45:37 +02:00
Pavel Emelyanov	420baf5035	api: Make get_highest_supported_sstable_version use proper service This endpoint now grabs one via database -> table -> sstables manager chain, but there's shorter route, namely via sstables format selector. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-28 10:18:57 +03:00
Pavel Emelyanov	61c8b571e5	api: Move system::get_highest_supported_sstable_version set/unset It's currently registered with all other system endpoints and is not unregistered. Its correct place is in the sstables-format-selector set/unset functions. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-28 10:18:23 +03:00
Pavel Emelyanov	f090bdabbb	api: Scaffold for sstables-format-selector This "service" will have its own endpoint soon Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-28 10:17:38 +03:00
Botond Dénes	31342ecb5d	Merge 'tasks: fix virtual tasks children' from Aleksandra Martyniuk Fix how regular tasks that have a virtual parent are created in task_manager::module::make_task: set sequence number of a task and subscribe to module's abort source. Fixes: #21278. Needs backport to 6.2 Closes scylladb/scylladb#21280 * github.com:scylladb/scylladb: tasks: fix sequence number assignment tasks: fix abort source subscription of virtual task's child	2024-10-28 08:59:40 +02:00
Aleksandra Martyniuk	85d9565158	test: repair: drop log checks from test_repair_succeeds_with_unitialized_bm Currently, test_repair_succeeds_with_unitialized_bm checks whether repair finishes successfully and the error is properly handled if batchlog_manager isn't initialized. Error handling depends on logs, making the test fragile to external conditions and flaky. Drop the error handling check, successful repair is a sufficient passing condition. Fixes: #21167. Closes scylladb/scylladb#21208	2024-10-28 08:39:16 +02:00
Botond Dénes	416159e5d9	Merge 'docs/alternator: explain service discovery HTTP requests' from Nadav Har'El Add a description of the service discovery HTTP requests - `/` and `/localnodes` that was previously not documented except in a design document that is unfortunately no longer available publically (https://docs.google.com/document/d/1twgrs6IM1B10BswMBUNqm7bwu5HCm47LOYE-Hdhuu_8/edit). Fixes https://github.com/scylladb/scylladb/issues/20989 Developer-oriented documentation so no need to backport. Closes scylladb/scylladb#21000 * github.com:scylladb/scylladb: docs/alternator: explain service discovery HTTP requests docs/alternator: split Alternator-specific APIs from alternator.md	2024-10-28 08:21:28 +02:00
Benny Halevy	2268912589	docs: add documentation for scylla_identifier Commit `3a12ad96c7` added an sstable_identifier uuid to the SSTable scylla_metadata component, however it was under-documented and this patch adds the missing documentation for the sstable component format, and to the scylla sstable tool documentation. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#21221	2024-10-28 08:18:08 +02:00
Kefu Chai	0f9d2ab577	build: cmake: disable Seastar exception hack in `cc3953e5`, we disabled Seastar exception hack in configure.py. this change disabled the Seastar exception hack in the following two builds: - build generated directly by configure.py - build configured with multi-config generator using CMake but we also have non-multi-config build using CMake. to be more consistent, let's apply the equivalent change to non-multi-config build of CMake. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21233	2024-10-28 08:11:43 +02:00
Botond Dénes	be70755f47	Merge 'repair: Fix finished ranges metrics for removenode' from Asias He The skipped ranges should be multiplied by the number of tables Otherwise the finished ranges ratio will not reach 100%. Fixes #21174 Closes scylladb/scylladb#21252 * github.com:scylladb/scylladb: test: Add test_node_ops_metrics.py repair: Make the ranges more consistent in the log repair: Fix finished ranges metrics for removenode	2024-10-28 08:09:32 +02:00
Asias He	9868ccbac0	test: Add test_node_ops_metrics.py It tests the node_ops_metrics_done metric reaches 100% when a node ops is done. Refs: #21174	2024-10-28 08:45:37 +08:00
Pavel Emelyanov	2f9f76fddf	sstables_loader: Mark to_replica_set() private It's not called from outside Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21210	2024-10-27 22:28:54 +02:00
Anna Stuchlik	ef4bcf8b3f	doc: remove the Cassandra references from notedool This PR removes the reference to Cassandra from the nodetool index, as the native nodetool is no longer a fork. In addition, it removes the Apache copyright. Fixes https://github.com/scylladb/scylladb/issues/21238 Closes scylladb/scylladb#21240	2024-10-27 22:26:33 +02:00
Kefu Chai	e65185ba6f	build: merge scylla's and seastar's compilation database Since commit `415c83fa`, Seastar is built as an external project. As a result, the compile_commands.json file generated by ScyllaDB's CMake build system no longer contains compilation rules for Seastar's object files. This limitation prevents tools from performing static analysis using the complete dependency tree of translation units. This change merges Seastar's compilation database with ScyllaDB's and places the combined database in the source root directory, maintaining backward compatibility. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21234	2024-10-27 22:01:29 +02:00
Tomasz Grabiec	850d9cfb59	node-exporter: Disable hwmon collector This collector reads nvme temperature sensor, which was observed to cause bad performance on Azure cloud following the reading of the sensor for ~6 seconds. During the event, we can see elevated system time (up to 30%) and softirq time. CPU utilization is high, with nvm_queue_rq taking several orders of magnitude more time than normally. There are signs of contention, we can see __pv_queued_spin_lock_slowpath in the perf profile, called. This manifests as latency spikes and potentially also throughput drop due to reduced CPU capacity. By default, the monitoring stack queries it once every 60s. Closes scylladb/scylladb#21165	2024-10-27 21:59:15 +02:00
Kefu Chai	f5b29331a2	build: populate --enable-dist --disable-dist to CMake before this change, the "dist" targets are always enabled in the CMake-based building system. but the build rules generated by `configure.py` does respect `--enable-dist` and `--disable-dist` command line options, and enable/distable the dist targets respectively. in this change, we - add an CMake option named "Scylla_DIST". the "dist" subdirectory in CMake only if this option is ON. - pouplate the `--enable-dist` and `--disable-dist` option down to cmake by setting the `Scylla_DIST` option, when creating the build system using CMake. this enables the CMake-based build system to be functionality wise more closer to the legacy building system. Refs scylladb/scylladb#2717 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21253	2024-10-27 21:57:46 +02:00
Kefu Chai	24d14b601b	treewide: s/boost::adaptors::map_values/std::views::values/ now that we are allowed to use C++23. we now have the luxury of using `std::views::values`. in this change, we: - replace `boost::adaptors::map_values` with `std::views::values` - update affected code to work with `std::views::values` - the places where we use `boost::join()` are not changed, because we cannot use `std::views::concat` yet. this helper is only available in C++26. to reduce the dependency to boost for better maintainability, and leverage standard library features for better long-term support. this change is part of our ongoing effort to modernize our codebase and reduce external dependencies where possible. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21265	2024-10-27 21:32:45 +02:00
Avi Kivity	3124711fc4	Merge 'Report rows_merged in compaction_history rest api and nodetool' from Łukasz Paszkowski Currently, running the `nodetool compactionhistory` command or using the rest api `curl -X GET --header "Accept: application/json" "http://localhost:10000/compaction_manager/compaction_history"` return compaction history without the `row_merged` field. The series computes rows merged during compaction and provides this information to users via both the nodetool command and the rest api. The `rows_merged` field contains information on merged clustering keys across multiple sstable files. For instance, compacting two sstables of a table consisting of 7 rows where two rows are part of the both sstables, the output would have the following format: {1: 5, 2: 2}. No backport is required. It extends the existing compaction history output. Fixes https://github.com/scylladb/scylladb/issues/666 Closes scylladb/scylladb#20481 * github.com:scylladb/scylladb: test/rest_api: Add tests for compactionhistory nodetool: Add rows merged stats into compactionhistory output compaction: Update compaction history with collected histogram compaction: Remove const qualifier from methods creating sstable readers sstable_set: Add optional statistics to make_local_shard_sstable_reader make_combined_reader: Add optional parameter, combined_reader_statistics reader_selector: Extend with maximum reader count mutation_fragment_merger: Create histogram while consuming mutation fragment batches	2024-10-27 21:26:11 +02:00
Kefu Chai	158008dd2c	mutation_writer: simplify classification using with_deserialized() return value Since `with_deserialized()` returns the lambda function's result, we can directly return the bucket from within the lambda instead of relying on side effects. This makes the code more explicit and functional. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21273	2024-10-27 21:20:55 +02:00
Nadav Har'El	6fdd0ebd3b	RBAC: confirm that unprivileged users can't read the roles table A worry was raised that an unprivileged user might be able to read the system.roles table - which contains the Alternator secret keys (and also CQL's hashed passwords). This patch adds tests that show that this worry is unjustified - and acts as a regression test to ensure it never becomes justified. The tests show that an unprivileged user cannot read the system.roles table using either CQL or Alternator APIs. More specifically, the two tests in this patch demonstrate that: * The Alternator API does not allow an unprivileged user to read ANY system table, unless explicitly granted permissions for that table. * The CQL API whitelists (see service::client_state::has_access) specific system tables - e.g., system_schema.tables - that are made readable to any unprivileged user. But the system.auth table is NOT whitelisted in this way - and is unreadable to unprivileged users unless explicitly granted permissions on that table. The new tests passes on both Scylla and Casssandra. Refs #5206 (that issue is about removing the Alternator secret keys from the roles table - but stealing CQL salted hashes is still pretty bad, so it's good to know that unprivileged users can't read them). Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#21215	2024-10-27 21:09:38 +02:00
Nadav Har'El	1634a64ffd	cql-pytest: test a few small materialized views CQL issues While documenting materialized view in a new document (Refs #16569) I encountered a few questions on how various CQL operations work on a table that has views, and this patch contains tests that clarify their answer - and can later guarantee that the answer doesn't unintentionally change in the future. The questions that these tests answer are: 1. That TRUNCATE on a base table also TRUNCATEs its views. This is just a basic test, with no attempt to reproduce issue #17635 (which is about the truncation of the base and views not being atomic). 2. That DROP TABLE is not allowed on a base table that has views. 3. That DROP KEYSPACE is allowed, even if there are tables with views. 4. Test that ALTER TABLE tbl DROP is never allowed in Cassandra, but allowed in some cases by Scylla 5. Test that ALTER TABLE tbl ADD is allowed, and "SELECT *" expands to select the new column into the materialized view as well. All the new tests pass on both Scylla and Cassandra. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#21142	2024-10-27 21:08:28 +02:00
Botond Dénes	7c75fc599f	streaming: stream-session: switch to tracking permit The stream-session is the receiving end of streaming, it reads the mutation fragment stream from an RPC stream and writes it onto the disk. As such, this part does no disk IO and therefore, using a permit with count resources is superfluous. Furthermore, after `d98708013c`, the count resources on this permit can cause a deadlock on the receiver end, via the `db::view::check_view_update_path()`, which wants to read the content of a system table and therefore has to obtain a permit of its own. Switch to a tracking-only permit, primarily to resolve the deadlock, but also because admission is not necessary for a read which does no IO. Refs: scylladb/scylladb#20885 (partial fix, solves only one of the deadlocks) Fixes: scylladb/scylladb#21264 Closes scylladb/scylladb#21059	2024-10-27 20:01:25 +02:00
Avi Kivity	7ffbfe8bb3	Merge 'Squash some sstables::test helpers' from Pavel Emelyanov There's a `missing_summary_first_last_sane` test case that uses some very specific way of modifying an sstable -- it loads one from resources, then tries to "write" the loaded stuff elsewhere. For that it uses a special purpose test::store() helper and a bunch of auxiliary ones from the same class. Those aux helpers are not used anywhere else and are also very special for this test case, so it make sense to keep this whole functionality in a single helper. Closes scylladb/scylladb#21255 * github.com:scylladb/scylladb: test: Squash test::change_generation_number() into test::store() test: Squash test::change_dir() into test::store() test: Coroutinize sstables::test::store()	2024-10-27 19:59:59 +02:00
Anna Stuchlik	aa0dadea48	doc: extend the ToC for CDC This commit adds the missing links to the CDC index page. Fixes https://github.com/scylladb/scylladb/issues/21137 Closes scylladb/scylladb#21286	2024-10-27 19:57:59 +02:00
Anna Stuchlik	b2b9622e32	doc: fix redundant references to version 6.2 This commit removes mentions of version 6.2 that were introduced with https://github.com/scylladb/scylladb/pull/17969. Now that the documentation is versioned, there should be no reference to specific versions. Fixes https://github.com/scylladb/scylladb/issues/21276 Closes scylladb/scylladb#21277	2024-10-27 14:47:40 +02:00
Paweł Zakrzewski	b077685fec	test/cql-pytest: GROUP BY with static columns This commit adds a new test case 'test_group_by_static_column_and_tombstones' to verify the behavior of GROUP BY queries with static columns. The test is adapted from Cassandra's test suite and aims to reproduce issue #21267. Original, larger test: cassandra_tests/validation/operations/select_group_by_test.py::testGroupByWithPaging() Closes scylladb/scylladb#21270	2024-10-27 14:45:53 +02:00
Aleksandra Martyniuk	910a6fc032	tasks: fix sequence number assignment Currently, children of virtual tasks do not have sequence number assigned. Fix it.	2024-10-25 15:30:13 +02:00
Aleksandra Martyniuk	1eb47b0bbf	tasks: fix abort source subscription of virtual task's child Currently, if a regular task does not have a parent or its parent is a virtual tasks then it subscribes to module's abort source in task_manager::task::impl constructor. However, at this point the kind of the task's parent isn't set. Due to that, children of virtual tasks aren't aborted on shutdown. Subscribe to module's abort source in task::impl::set_virtual_parent.	2024-10-25 14:18:00 +02:00
Kefu Chai	e7d6ab576b	backup_task: remove unused member variable Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21258	2024-10-25 11:49:06 +03:00
Abhinav	c00d40b239	fix nodetool status to show zero-token nodes In the current scenario, the nodetool status doesn’t display information regarding zero token nodes. For example, if 5 nodes are spun by the administrator, out of which, 2 nodes are zero token nodes, then nodetool status only shows information regarding the 3 non-zero token nodes. This commit intends to fix this issue by leveraging the “/storage_service/host_id ” API and adding appropriate logic in scylla-nodetool.cc to support zero token nodes. Robust topology tests are added, which spins up scylla nodes and confirm nodetool status output for various cases, providing good coverage. A test is also added in nodetool/test_status.py to verify this logic. These tests fail without this commit’s zero token node support logic, hence verifying the behavior. The test `test_status_keyspace_joining_node` has been removed. This test is based on case where host_id=None, which is impossible. Since we now use host_id_map for node discovery in nodetool, the nodes with "host_id=None" go undetected. Since this case is anyway impossible, we can get rid of this. This PR fixes a bug. Hence we need to backport it. Backporting needs to be done only to 6.2 version, since earlier versions dont support zero token nodes. Fixes: scylladb/scylladb#19849	2024-10-25 13:28:09 +05:30
Abhinav	39dfd2d7ac	test: move `wait_for_first_completed` to pylib/util.py This function is needed in a new test added in the next commit and this refactoring avoids code duplication.	2024-10-25 13:26:42 +05:30
Abhinav	72f3c95a63	token_metadata: rename endpoint_to_host_id_map getter and add support for joining nodes Rename host_id map getter, 'get_endpoint_to_host_id_map_for_reading' to 'get_endpoint_to_host_id_map_' Also modify the getter to return information regarding joining nodes as well. This getter will later be used for retrieving the nodes in nodetool status, hence it needs to show all nodes, including joining ones. The function name suffix `_for_reading` suggests that the function was used in some other places in the past, and indeed if we need endpoints "for reading" then we cannot show joining endpoints. But it was confirmed that this function is currently only used by "/storage_service/host_id" endpoint, hence it can be modified as required. Fixes: scylladb/scylladb#17857	2024-10-25 13:20:27 +05:30
Pavel Emelyanov	5e713b2b14	Merge 'dht: remove unused #includes ' from Kefu Chai these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#21237 * github.com:scylladb/scylladb: .github: add dht to iwyu's CLEANER_DIR dht: remove unused `#include`s	2024-10-24 18:40:49 +03:00
Pavel Emelyanov	7595ef7303	test: Squash test::change_generation_number() into test::store() No other usages of the former helper other than immediatelly followed by the latter, no point in keepint it around. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-24 11:29:17 +03:00
Pavel Emelyanov	e885b0e6cd	test: Squash test::change_dir() into test::store() No other usages of the former helper other than immediatelly followed by the latter, no point in keepint it around. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-24 11:28:39 +03:00
Pavel Emelyanov	874cf2ea6f	test: Coroutinize sstables::test::store() Ahead of future changes Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-24 11:28:07 +03:00
Benny Halevy	5498018cbe	data_dictionary: keyspace_metadata::describe: print tablets enabled also when defaulted Now that tablets may be explicitly enabled when creating a new keyspace, describe tablets as enabled even when the default initial_tablets==0 is used. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-24 10:18:42 +03:00
Benny Halevy	63cbb6e071	tablets_test: test enable/disable tablets when creating a new keyspace Test both configuration values for `enable_tablets` and the possibility to explicitly enable or disable tablets, respectively, when creating a keyspace using the `tablets = {'enabled': true\|false}` CREATE KEYSPACE option. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-24 10:18:42 +03:00
Benny Halevy	b0e12cb40d	treewide: always allow tablets keyspaces With the tablets feature always enabled (Unless gossip toopology changes are forced), the enable_tablets option now controls only the default for newly created keyspaces. Even when set to `false`, tablets are still enabled as a feature and the user may explicitly enable tablets using `CREATE KEYSPACE <name> WITH tablets = {'enabled': true}` Note: best viewed with `git show -w` Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-24 10:18:42 +03:00
Benny Halevy	bc62407421	feature_service: prevent enabling both tablets and gossip topology changes Tablets require raft consistent topology changes. Therefore, document that they are incompatible in the config help and prevent their usage in `feature_config_from_db_config` Fixes scylladb/scylladb#21075 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-24 10:18:42 +03:00
Benny Halevy	9ef2dc2428	alternator: create_keyspace_metadata: enable tablets using feature_service Rather than using the local configuration option on this node, check the cluster feature instead. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-24 10:18:42 +03:00
Asias He	1392a6068d	repair: Make the ranges more consistent in the log Consider the number of tables for the number of ranges logging. Make it more consistent with the log when the ops starts.	2024-10-24 10:31:15 +08:00
Asias He	cffe3dc49f	repair: Fix finished ranges metrics for removenode The skipped ranges should be multiplied by the number of tables. Otherwise the finished ranges ratio will not reach 100%. Fixes #21174	2024-10-24 10:31:15 +08:00
Kefu Chai	a9e18f70b0	Revert submodule change in `6ead5a4696` in `6ead5a46`, we included submodule changes in cqlsh and java by accident. this was not intended. and this broke the artifacts-rocky8-test. in this change, both changes in the submodule are reverted. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21236	2024-10-23 19:49:20 +03:00
Pavel Emelyanov	9014da26e1	Merge 'docs: reference object storage config doc from nodetool commands ' from Kefu Chai this series: - promote object storage configuration to user-facing documentation - reference object storage config doc from nodetool commands --- the nodetool backup/restore commands are not included by any LTS branches yet, hence no need to backport. Closes scylladb/scylladb#21071 * github.com:scylladb/scylladb: docs: move keyspace-storage-option from cql-extensions to admin docs: reference admin.rst for object storage config docs: reference object storage config doc from nodetool commands docs: promote object storage configuration to user-facing documentation	2024-10-23 19:41:46 +03:00
Michał Jadwiszczak	68d0c9a18a	test/auth_cluster/test_raft_service_levels: match enterprise SL limit Despite OSS doesn't limit number of created service levels, match the enterprise limit to decrease divergence in the test between OSS and enterprise. Fixes scylladb/scylladb#21044 Closes scylladb/scylladb#21045	2024-10-23 17:44:19 +02:00
Kefu Chai	bea18f0571	.github: add dht to iwyu's CLEANER_DIR to avoid future violations of include-what-you-use. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-23 17:45:14 +08:00
Kefu Chai	8d1b3223ab	dht: remove unused `#include`s these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-23 17:45:14 +08:00
Dawid Mędrek	298cafff35	cql-pytest/test_describe: Introduce auxiliary type for service levels We introduce an auxiliary type representing a service level for making it easier to adjust the tests in Enterprise. We move the responsibility of producing create statements for service levels to the class, so we only need to modify the code in one place when necessary. All existing relevant tests have been adjusted to this change. Closes scylladb/scylladb#21230	2024-10-23 10:15:25 +02:00
Kamil Braun	f5c60e538d	Merge 'cql/tablets: fix retrying ALTER tablets KEYSPACE' from Piotr Smaron ALTER tablets-enabled KEYSPACES (KS) may fail due to `group0_concurrent_modification`, in which case it's repeated by a `for` loop surrounding the code. But because raft's `add_entry` consumes the raft's guard (by `std::move`'ing the guard object), retries of ALTER KS will use a moved-from guard object, which is UB, potentially a crash. The fix is to remove the before mentioned `for` loop altogether and rethrow the exception, as the `rf_change` event will be repeated by the topology state machine if it receives the concurrent modification exception, because the event will remain present in the global requests queue, hence it's going to be executed as the very next event. Note: refactor is implemented in the follow-up commit. Fixes: scylladb/scylladb#21102 Should be backported to every 6.x branch, as it may lead to a crash. Closes scylladb/scylladb#21121 * github.com:scylladb/scylladb: test: add UT to test retrying ALTER tablets KEYSPACE cql/tablets: fix indentation in `rf_change` event handler cql/tablets: fix retrying ALTER tablets KEYSPACE	2024-10-23 10:01:21 +02:00
Botond Dénes	519e167611	Merge 'replica/table: check memtable before discarding tombstone during read' from Lakshmi Narayanan Sreethar On the read path, the compacting reader is applied only to the sstable reader. This can cause an expired tombstone from an sstable to be purged from the request before it has a chance to merge with deleted data in the memtable leading to data resurrection. Fix this by checking the memtables before deciding to purge tombstones from the request on the read path. A tombstone will not be purged if a key exists in any of the table's memtables with a minimum live timestamp that is lower than the maximum purgeable timestamp. Fixes #20916 `perf-simple-query` stats before and after this fix : `build/Dev/scylla perf-simple-query --smp=1 --flush` : ``` // Before this Fix // --------------- 94941.79 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59393 insns/op, 24029 cycles/op, 0 errors) 97551.14 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59376 insns/op, 23966 cycles/op, 0 errors) 96599.92 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59367 insns/op, 23998 cycles/op, 0 errors) 97774.91 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59370 insns/op, 23968 cycles/op, 0 errors) 97796.13 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59368 insns/op, 23947 cycles/op, 0 errors) throughput: mean=96932.78 standard-deviation=1215.71 median=97551.14 median-absolute-deviation=842.13 maximum=97796.13 minimum=94941.79 instructions_per_op: mean=59374.78 standard-deviation=10.78 median=59369.59 median-absolute-deviation=6.36 maximum=59393.12 minimum=59367.02 cpu_cycles_per_op: mean=23981.67 standard-deviation=32.29 median=23967.76 median-absolute-deviation=16.33 maximum=24029.38 minimum=23947.19 // After this Fix // -------------- 95313.53 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59392 insns/op, 24058 cycles/op, 0 errors) 97311.48 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59375 insns/op, 24005 cycles/op, 0 errors) 98043.10 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59381 insns/op, 23941 cycles/op, 0 errors) 96750.31 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59396 insns/op, 24025 cycles/op, 0 errors) 93381.21 tps ( 71.1 allocs/op, 0.0 logallocs/op, 14.1 tasks/op, 59390 insns/op, 24097 cycles/op, 0 errors) throughput: mean=96159.93 standard-deviation=1847.88 median=96750.31 median-absolute-deviation=1151.55 maximum=98043.10 minimum=93381.21 instructions_per_op: mean=59386.60 standard-deviation=8.78 median=59389.55 median-absolute-deviation=6.02 maximum=59396.40 minimum=59374.73 cpu_cycles_per_op: mean=24025.13 standard-deviation=58.39 median=24025.17 median-absolute-deviation=32.67 maximum=24096.66 minimum=23941.22 ``` This PR fixes a regression introduced in `ce96b472d3` and should be backported to older versions. Closes scylladb/scylladb#20985 * github.com:scylladb/scylladb: topology-custom: add test to verify tombstone gc in read path replica/table: check memtable before discarding tombstone during read compaction_group: track maximum timestamp across all sstables	2024-10-23 10:28:00 +03:00
Botond Dénes	d6a79fefda	Merge 'Do not leak S3 file-uploading parts on exceptions' from Pavel Emelyanov File uploading code spawns all parts uploading into background. If this "spawning" fails (not the uploading code itself), any fiber that was spawned before is orphaned. It will eventually stop on its own, by while it's alive it may use(-after-free) the do_upload_file object. Another issue with not handling spawn exception, is that multipart upload object is not aborted in this case. So it's leaked until garbage collector picks it up, which is not critical, but unpleasant. Closes scylladb/scylladb#21139 * github.com:scylladb/scylladb: s3/client: Restore indentation after previous patch s3/client: Catch do_upload_file::upload_part() exceptions	2024-10-23 10:12:29 +03:00
Ernest Zaslavsky	59e2ed884d	Update seastar submodule * seastar abd20efd...f821bda1 (17): > http: http status classification > loopback: add pending capacity param and fix deadlock in httpd_test > allow setting buffer sizes on server_socket > core: add missing assert header to chunked_fifo > cmake: Don't emit message when searching for libarchive > stall-analyser: pass args.tmin instead of tmin > build: do not check for CMAKE_CXX_STANDARD < 20 > README.md: specify CMAKE_CXX_STANDARD in the sample > cmake: Fix DPDK libarchive dep > c-ares: update cooking version to 1.32.3 > build: support c-ares >= 1.34.1 > iotune: clarify fsqual error message > Make total_steal_time() monotonic. > Remove account_idle > reactor: add better sleep time accounting > reactor: add cpu and awake time reactor metrics > Zero-init total sleep time Closes scylladb/scylladb#21225	2024-10-23 09:30:56 +03:00
Botond Dénes	b9b778054a	Merge 'test.py: Add option to fail after number of failures' from Petr Hála * Add `--max-failures` flag to test.py, which will stop the execution after number of failures * Helps with "fails-fast" approach and can be used to improve CI speed, especially the 100times run * Adds the number of cancelled tests to both summary and junit xml. I did not include them in boost, since it does not contain any statistics. * Removes unnecessary list creation in test.py * Completely unrelated change, but it is small enough that I feel it can be included as part of this one. If this is an issue I can create separate PR for it * Add `Test.started` property * Helps with determining the current status of the Test and differentiating cancelled/not started tests. * Add `Test.failed` and `Test.did_not_run` read-only computed properties * Helper methods to determine status, instead of using `Test.success`, which does not tell the entire story * Fix `ScyllaClusterManager.stop()` method, so it doesn't fail when ran multiple times * This happens when tasks are cancelled, not sure yet why, it almost certainly non-wanted behaviour but this behaviour was already there and with this fix it no longer causes errors I will use backport/None for now as it is a new feature. Fixes https://github.com/scylladb/qa-tasks/issues/1714 Closes scylladb/scylladb#21098 * github.com:scylladb/scylladb: test.py: Add option to fail after number of failures test.py: Add started, failed and did_not_run properties to Test test.py: Remove unnecessary list creation test: lib: Fix ScyllaClusterManager.stop()	2024-10-23 09:11:52 +03:00
Kefu Chai	6a7eaea9f4	mutation_writer/feed_writer: remove redundant check `mutation_reader::is_end_of_stream()` returns `_impl->is_end_of_stream() && is_buffer_empty()`, so `!is_end_of_stream()` equals to ` `!_impl->is_end_of_stream() \|\| !is_buffer_empty()`, which in turn always equals to `!_impl->is_end_of_stream() \|\| !is_buffer_empty() \|\| !is_buffer_empty()`. hence there is no need to check `rd.is_buffer_empty()` again. in this change, the redundant condition is dropped. simpler this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21224	2024-10-23 08:48:08 +03:00
Avi Kivity	cc3953e504	build: disable Seastar exception hack In [1], Seastar started to bypass a lock in libgcc's exception throwing mechanism to allow scalability on large machines. The problem is documented in [2] and reported as fixed. In [3], testing results on a 2s96c192t machine are reported. The problem appears indeed fixed with gcc 14's runtime (which we use, even though we build with clang). Given the new results, we can safely drop the exception scalability hack. As [1] states that the hack causes the loss of a translation cache, we may gain some performance this way. With that, we disable the cache by defining some random macro. [1] https://github.com/scylladb/seastar/464f5e3ae43b366b05573018fc46321863bf2fae [2] https://gcc.gnu.org/bugzilla/show_bug.cgi?id=71744 [3] https://github.com/scylladb/seastar/issues/2479#issuecomment-2427098413 Closes scylladb/scylladb#21217	2024-10-22 22:20:07 +03:00
Nadav Har'El	5fd3177057	Merge 'mv: add a dedicated read concurrency semaphore for view update read before writes' from Wojciech Mitros When writing to some tables with materialized views, we need to read from the base table first to perform a delete of the old view row. When doing so, the memory used for the read is tracked by the user read concurrency semaphore. When we have a large number of such reads, we may use up all of the semaphore units, causing the following reads to be queued. When we have some user reads coming at the same time, these reads can have very high latency due to the write workload on the base table. We want to avoid this, so that the write workload doesn't have a high impact on the latency of the read workload. This is fixed in this patch by adding a separate read concurrency semaphore just for view update read-before-writes. With the new semaphore, even if there are many view update read-before-writes, they will be queued on a different semaphore than the user reads, and they won't impact their latency. The second issue fixed by this patch is the concurrency of the view updates that is currently unlimited. Because of that view updates may take up so much memory that they we may run out of memory. This is fixed by using the read admission on the view update concurrency semaphore. This limits the number of concurrent view update reads to max_count_concurrent_view_update_reads, all other incoming view update reads are queued using just a small chunk of memory. Without this, the reads would also get queued after exceeding view_update_reader_concurrency_semaphore_serialize_limit_multiplier, but they would take much more memory while staying in the queue. The new semaphore has half the capacity of the regular user read concurrency semahpore and is currently used only for user writes - is't used independently of the scheduling group on which we base the read semaphore selection, but we use a different code path for streaming (not database::do_apply) and we shouldn't have view updates in system writes or during compaction. This patch also adds a test to confirm that the view update workload doesn't impact the read latency, as well as a test which confirms that we do not run out of memory even under heavy view udpate workload. The issue of view updates causing increased latencies most often occurs in the following scenario: * we have a medium to high write workload to a table with a materialized view which requires reading from the base table before sending the update to delete the old rows * we have any read workload * one replica is slower or is handling more writes due to an imbalance of data distribution * we write with a cl<ALL, the mentioned replica is replying to write requests slower while new ones keep being sent to it. * each write performs a read first taking resources from the user read concurrency semaphore, so when enough writes accumulate the reads using the semaphore start getting queued * the queue is shared by regular reads and view update reads. When there's enough view update reads in the queue, regular reads start getting increased latencies An sct test (perf-regression-latency-mv-read-concurrency) was prepared to somewhat resemble this scenario: * the tables were prepared satisfying the conditions above * we use a medium write workload and a very low read workload * the imbalance is achieved by writing to just a few (10) partitions - some replicas (and shards) can have twice or more used partitions than others. We also keep writing to a limited (though high) number of rows, to cause overwrites which require reading before sending the view update * to minimize the test case, we use a cluster of 3 nodes and rf=2, we write with cl=ONE to have background replica writes and read with cl=ALL to wait for the slower replica to respond. In the test above: * without the fix, the latency of reads increases over 50s * with the fix, the latency of reads stays below 20ms Fixes https://github.com/scylladb/scylladb/issues/8873 Fixes https://github.com/scylladb/scylladb/issues/15805 The patch is not that small and it isn't fixing a regression, so no backports Closes scylladb/scylladb#20887 * github.com:scylladb/scylladb: test: add test for high view update concurrency causing bad_allocs test: add test for high view update concurrency degrading read latency mv: add a dedicated read concurrency semaphore for view update read before writes	2024-10-22 22:17:23 +03:00
Aleksandra Martyniuk	878a12c922	test: change quotation marks Before python 3.12 formatted strings couldn't have reused quotes. Change the type of quotation mark in get_cgroup so it could be used with earlier python versions. Closes scylladb/scylladb#21209	2024-10-22 20:42:05 +03:00
Piotr Smaron	522bede8ec	test: add UT to test retrying ALTER tablets KEYSPACE The newly added testcase is based on the already existing `test_alter_dropped_tablets_keyspace`. A new error injection is created, which stops the ALTER execution just before the changes are submitted to RAFT. In the meantime, a new schema change is performed using the 2nd node in the cluster, thus causing the 1st node to retry the ALTER statement.	2024-10-22 18:22:01 +02:00
Piotr Smaron	3f4c8a30e3	cql/tablets: fix indentation in `rf_change` event handler Just moved the code that previously was under a `for` loop by 1 tab, i.e. 4 spaces, to the left.	2024-10-22 18:22:01 +02:00
Piotr Smaron	de511f56ac	cql/tablets: fix retrying ALTER tablets KEYSPACE ALTER tablets-enabled KEYSPACES (KS) may fail due to `group0_concurrent_modification`, in which case it's repeated by a `for` loop surrounding the code. But because raft's `add_entry` consumes the raft's guard (by `std::move`'ing the guard object), retries of ALTER KS will use a moved-from guard object, which is UB, potentially a crash. The fix is to remove the before mentioned `for` loop altogether and rethrow the exception, as the `rf_change` event will be repeated by the topology state machine if it receives the concurrent modification exception, because the event will remain present in the global requests queue, hence it's going to be executed as the very next event. `topology_coordinator::handle_topology_coordinator_error` handling the case of `group0_concurrent_modification` has been extended with logging in order not to write catch-log-throw boilerplate. Note: refactor is implemented in the follow-up commit. Fixes: scylladb/scylladb#21102	2024-10-22 18:22:00 +02:00
Avi Kivity	ec543e3902	Merge 'Remove all_datadirs vector of strings from table::config' from Pavel Emelyanov The all_datadirs keeps paths to directories where local sstables can be. In fact, Scylla doesn't put sstables there, but can try to find them on boot and when checking snapshots. The 0th element of this vector, called datadir, had recently been removed by #20675, now it's time to drop all_datadirs as well. The needed paths can be obtained from table's storage options (see #20542) and db::config::data_file_directories option. Closes scylladb/scylladb#21212 * github.com:scylladb/scylladb: sstables: Open-code format_table_directory_name() moved recently replica,sstables: Move format_table_directory_name() table: Remove all_datadirs sstables: Generate table::all_datadirs from db::config and storage_options replica: Prepare vector of fs::path-s with table dirs table: Check storage options in get_snapshot_details()	2024-10-22 17:21:31 +03:00
Laszlo Ersek	63417f6a57	utils/small_vector: refactor expansion condition in reserve*() Rewrite _begin + n > _capacity_end as n > _capacity_end - _begin and then as n > capacity() for two reasons: - The last form is easier to read than the first form. - Per N4950 (the final C++23 working draft), [expr.add] paragraph 4, the expression _begin + n (i.e., P + J) is defined only if 0 ≤ 0 + n ≤ _capacity_end - _begin (i.e., 0 ≤ i + j ≤ n) equivalently, only if _begin ≤ _begin + n ≤ _capacity_end Therefore, the expression _begin + n invokes undefined behavior exactly when we'd expect our check _begin + n > _capacity_end to evaluate to true. gcc and clang have been aggressively equating undefined behavior to "never happens"; let's prevent that here. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#21213	2024-10-22 17:12:11 +03:00
Avi Kivity	847c850034	schema: add accessors for primary key columns and non-primary-key columns It's somewhat common to ask for the partition key and clustering key columns, or for the static and regular columsn. Provide accessors for them rather than requiring the user to glue them. Some callers are converted. Closes scylladb/scylladb#21191	2024-10-22 15:01:14 +02:00
pehala	870f3b00fc	test.py: Add option to fail after number of failures Add --max-failures configuration option to specify the amount, if not set, or not positive, it will never trigger. Update also the junit reporting to include skipped tests	2024-10-22 13:29:34 +02:00
pehala	c1dd97a049	test.py: Add started, failed and did_not_run properties to Test This ensures we can determine where in the execution pipeline the test currently is. failed and did_not_run are helper properties	2024-10-22 13:29:19 +02:00
pehala	e34dec71e7	test.py: Remove unnecessary list creation Using generators & set constructor, we can get rid of unnecessary list creation	2024-10-22 13:29:18 +02:00
pehala	16cd3fccdd	test: lib: Fix ScyllaClusterManager.stop() When cancelling running tasks, stop() could run multiple times and fail. Removed usage of del and added checks to ensure it won't crash.	2024-10-22 13:29:18 +02:00
Kefu Chai	7a1e067b4e	docs: move keyspace-storage-option from cql-extensions to admin as the admin needs to known the name of the experimental feature option they need to enable. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-22 18:30:29 +08:00
Kefu Chai	6f97c86a2b	docs: reference admin.rst for object storage config instead of repeating it in cql-extensions.md, let's reference the object storage related settings in admin.rst Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-22 18:26:19 +08:00
Kefu Chai	fe13b4e10e	docs: reference object storage config doc from nodetool commands Enhance the documentation for nodetool commands that use the `--endpoint` option by linking to the object storage configuration guide. This change provides users with essential context and detailed setup instructions for S3-compatible storage endpoints. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-22 18:26:19 +08:00
Kefu Chai	9bd9ee9f36	docs: promote object storage configuration to user-facing documentation this commit moves the object storage configuration guide from the developer documentation to the user-facing admin documentation. the change reflects the increasing importance of object storage integration in user-facing features. in this change: - move relevant content from `docs/dev/object_storage.md` to `docs/operating-scylla/admin.rst` - reformat the content from Markdown to reStructuredText (RST) - reword and restructure the content to be more user-friendly - add explanations and context suitable for a broader audience this change makes the object storage configuration information more accessible to Scylla administrators and end-users, supporting the adoption of new features built on top of object storage integration. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-22 18:26:19 +08:00
Benny Halevy	04d741bcbb	storage_service: on_change: update_peer_info only if peer info changed Return an optional peer_info from get_peer_info_for_update when the `app_state_map` arg does not change peer_info, so that we can skip calling update_peer_info, if it didn't change. Fixes scylladb/scylladb#20991 Refs scylladb/scylladb#16376 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#21152	2024-10-22 10:26:08 +02:00
Dawid Medrek	4ec0a014e3	docs/hinted-handoff: Add link to API reference We add a link to the API reference for the convenience of the user. Closes scylladb/scylladb#20065	2024-10-22 09:24:14 +03:00
pehala	28aa57f836	test.py: Refactor retry() Instead of metamethod that looks at all subclasses, use OOP with super() calls Closes scylladb/scylladb#21155	2024-10-22 09:23:30 +03:00
David Garcia	6b7b4addf9	docs: add dark theme to api Closes scylladb/scylladb#21161	2024-10-22 09:22:32 +03:00
pehala	59eb4eb528	test.py: Enhance progress report * Do not leave passed tests in between failed ones. * Use ANSI Escape sequences for manipulating console * Simplifies code and removes need for two object parameters Closes scylladb/scylladb#21176	2024-10-22 09:22:08 +03:00
Łukasz Paszkowski	34c05cb94f	test/rest_api: Add tests for compactionhistory For a table with NullCompactionStrategy and TimeWindowCompactionStrategy, the test - inserts a bunch of data and flushes the table - deletes/update some data, delete a range of data and flushes the table - Triggers a major compaction and calls for compactionhistory to retrieve and validate the histogram	2024-10-22 08:15:02 +02:00
Łukasz Paszkowski	8188a71787	nodetool: Add rows merged stats into compactionhistory output Incorporate rows merged statistics into the output of the compactionhistory command. Depending on the requested format type, the output has different form. For instance, compacting two sstables of a table consisting of 7 rows where two rows are part of the both sstables, the output would have the following format: text: {1: 5, 2: 2} json: [{"key":1,"value":5},{"key":2,"value":1}]} yaml: - key: 1 value: 5 - key: 2 value: 1	2024-10-22 08:15:02 +02:00
Łukasz Paszkowski	c01a38f3cf	compaction: Update compaction history with collected histogram A new field has been added to the compaction_stats structure to hold collected combined reader statistics. The struct is than used to update the compaction_history table.	2024-10-22 08:15:02 +02:00
Łukasz Paszkowski	7eac89da73	compaction: Remove const qualifier from methods creating sstable readers Compaction classes start mutate their internal members to be used in methods setup_sstable_reader and make_sstable_reader creating sstable reades that are marked as const. Remove the const qualifier from these methods. Even though it made sense initially to mark them as const, it is no longer applicable.	2024-10-22 08:15:02 +02:00
Łukasz Paszkowski	484655bf0d	sstable_set: Add optional statistics to make_local_shard_sstable_reader The pointer to combined_reader_statistics is propagated down to make_combined_reader in order to collect statistics. By default, a null pointer is propagated. Note that in case the pointer is valid and the sstable_set consists of exactly one sstable, statistics are skipped as all rows originate from exactly a single sstable file. The existing optimization is crucial `f75154afca`	2024-10-22 08:15:02 +02:00
Łukasz Paszkowski	a9f776494c	make_combined_reader: Add optional parameter, combined_reader_statistics All the overloaded make_combined_reader functions accept an optional pointer to combined_reader_statistics, to be propagated down through merging_reader to mutation_fragment_merger. By default, a null pointer is propagated.	2024-10-22 08:15:02 +02:00
Łukasz Paszkowski	84912c3155	reader_selector: Extend with maximum reader count The maximum reader count allows to predict the number of readers that can be created with create_new_readers(). This helps to correctly allocate a vector size in the rows_merged statistics when a combiner reader is created via make_combined_reader.	2024-10-22 08:15:02 +02:00
Łukasz Paszkowski	92f5c56afc	mutation_fragment_merger: Create histogram while consuming mutation fragment batches The mutation_fragment_merger takes one additional parameter in its constructor, that is a pointer to a combined_reader_statistics used to collect various statistics. The histogram is populated with data while the merger consumes batches from the producer and merges them into seperate mutation fragments. The size of the batch, that represents the number of streams the mutation fragment originates from, is used as a key in the historgam and its corresponding value is increased by one.	2024-10-22 08:15:02 +02:00
Botond Dénes	41de340d93	Merge 'Update get_description.py script' from Amnon Heiman get_description.py script is a document related script that looks for metrics description in the code. Its configuration needs to address changes in the code. This series contains a configuration change and a code fix that allows it to run as a standalone script, and not as a library. No need to backport, this a documentation related script. Closes scylladb/scylladb#19950 * github.com:scylladb/scylladb: scripts/get_description.py: param_mapping was missing scripts/metrics-config.yml: no need to get metrics from the tests	2024-10-22 08:42:15 +03:00
Kefu Chai	27fb893d9b	docs: nodetools-commands/restore: update to reflect the latest implementation in `787ea4b1d4`, we added "sstables" argument to the "nodetool restore" command. but we failed to update the document to reflect the change. in this change, we update the document for "restore" command to reflect the latest implementation changes introduced in commit `787ea4b1d4`: * Add information about the new "sstables" argument * Update command line usage of "--table" argument -- it is now madatory * Update the example accordingly. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21135	2024-10-22 08:30:06 +03:00
Kefu Chai	ce0a86c585	build: cmake: correct some tests' KIND before this change, we build some tests as if they are Seastar tests. but after `415c83fa`, these tests failed to link. because the Seastar::seastar_testing does not expose `-DSEASTAR_TESTING_MAIN` in its cflags. the behavior of the Seastar::seastar_testing is expected. because a test linking against this library is not necessarily driven by the `main()` provided by `testing/seastar_test.hh`. so, in this change, we correct the `KIND` parameter of these tests, so that they use `KIND BOOST`, as these tests can be driven by the `main()` provided by Boost.Test's driver. also there are some tests driven by Boost.Test's `main()`, but in the meanwhile, they utilize seastar_testing, so let's add `Seastar::seastar_testing` to their `LIBRARIES`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21183	2024-10-22 07:10:47 +03:00
Kefu Chai	6ead5a4696	treewide: move log.hh into utils/log.hh the log.hh under the root of the tree was created keep the backward compatibility when seastar was extracted into a separate library. so log.hh should belong to `utils` directory, as it is based solely on seastar, and can be used all subsystems. in this change, we move log.hh into utils/log.hh to that it is more modularized. and this also improves the readability, when one see `#include "utils/log.hh"`, it is obvious that this source file needs the logging system, instead of its own log facility -- please note, we do have two other `log.hh` in the tree. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-22 06:54:46 +03:00
Kefu Chai	6645cdf3b6	build: cmake: improve source generated for check_header in check_headers.cmake, we verify the self containness of a header file by replicating it and remove `#pragma once` directive in this header. but this approach failed to compile headers which include a header file with the same name in the root source directory, as we add `-I<directory-of-original-header>` in the cflags when building the generated source file, so that it can include the headers in the same directory. but this confuses the compiler, as, assuming we have "log.hh" in current directory, and under the root source directory, the compiler would always include the "log.hh" in the current directory even it should have included "log.hh" under the root source directory. in this change, instead of adding `-I<directory-of-original-header>` to cflags, we just include the header under test in a new .cc file solely generated for testing. this should address this problem. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21216	2024-10-22 06:28:16 +03:00
Kefu Chai	2d6af2791e	compaction: simplify time_window_compaction_strategy::get_window_lower_bound() since chrono allows dividion between durations with different units. let use it instead for rounding down to the nearest multiple of the window size, for better readability. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20476	2024-10-21 16:01:15 +03:00
Pavel Emelyanov	516a5f06a8	sstables: Open-code format_table_directory_name() moved recently This helper is small enough and it's easier to understand how table directory name is formatted without it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-21 15:18:19 +03:00
Pavel Emelyanov	eeb0d637bb	replica,sstables: Move format_table_directory_name() Now this helper is not needed in replica code, as all manipulations of tables' sstables now sit in the sstables/storage.cc. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-21 15:17:30 +03:00
Pavel Emelyanov	74728d3889	table: Remove all_datadirs It's write-only now, all the places than wanted to know where table's storage is (well -- "are", there can be several directories) already use storage_options. This finishes the work started by `9fe64b5d70`. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-21 15:15:54 +03:00
Pavel Emelyanov	dedb9d349c	sstables: Generate table::all_datadirs from db::config and storage_options As mentioned in the previous patch, there are several places that need to scan all datafile directories for a given table. This list is currently stored on table.config.all_datadirs, this patch stops using one and instead generates it from db::config::data_file_directories and table's storage options. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-21 15:13:27 +03:00
Pavel Emelyanov	0358515118	replica: Prepare vector of fs::path-s with table dirs Most of the time table with local storage keeps its sstables in a single directory referenced by its storage_options::local.dir path. However, there are two cases when code needs to check all datafile directories that could be configured -- on boot when distributed loader loads sstables, and when checking table snapshots. Both those places check table.cfg.all_datadirs vector of strings and convert strings to fs::path-s along the way. This patch prepares the vector of fs::path-s in advance and updates the loop code to work with path-s. This is preparation to next patching that will generate vector of paths for a table. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-21 15:10:53 +03:00
Pavel Emelyanov	4e329ba08f	table: Check storage options in get_snapshot_details() This is continuation of `24589cf00c` and `a734fd5c9c` -- if table is not based on local storage, getting snapshot details makes no sense. Another goal this change pursuits is to have storage_options::local object at hand to be used later. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-21 15:08:54 +03:00
Wojciech Mitros	4d719bacca	test: add test for high view update concurrency causing bad_allocs This commit add a test for checking whether a large view update workload can cause Scylla to run out of memory. In the test, we keep writing to a table table with a materialized view with a limited number of rows, causing overwrites which require reading from the table to perform view updates. Currently, due to the unlimited concurrency of view update reads, we may use too much memory which can lead to bad_allocs, causing Scylla to fail. To reach the failing state more consistently, we use add a sleep after reading the old value of the base row, to keep the reader concurrency semaphore units longer. At the same time, we use high concurrency and large row size to use up all Scylla's memory quickly. The test fails if Scylla runs out of memory and aborts, and succeeds otherwise.	2024-10-21 12:35:20 +02:00
Wojciech Mitros	f2c740710c	test: add test for high view update concurrency degrading read latency This commit add a test for checking whether a large view update workload impacts the latency of other user reads. In the test, we first create a table for reads and another table with a materialized view. We then start writing to the table with the view with a limited number of rows - when overwriting, we need to read the previous value of the row to prepare a delete of the old row in the view. This should not impact the latency of the read workload from the other table that we start at the same time. The test fails if any of the reads times out. To reach the failing state more consistantly, we use add a sleep after reading the old value of the base row, to keep the reader concurrency semaphore units longer. At the same time, we use a lower threshold for queueing reads on the semaphore, to see the impact of view update reads earlier. Because of the high load, the writes may timeout, but that's expected - we fail the test only if the user reads time out.	2024-10-21 12:34:55 +02:00
Kefu Chai	5cd619a60c	treewide: s/boost::adaptors::map_keys/std::views::keys/ now that we are allowed to use C++23. we now have the luxury of using `std::views::keys`. in this change, we: - replace `boost::adaptors::map_keys` with `std::views::keys` - update affected code to work with `std::views::keys` to reduce the dependency to boost for better maintainability, and leverage standard library features for better long-term support. this change is part of our ongoing effort to modernize our codebase and reduce external dependencies where possible. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21198	2024-10-21 12:47:52 +03:00
Wojciech Mitros	242079d70b	mv: add a dedicated read concurrency semaphore for view update read before writes When writing to some tables with materialized views, we need to read from the base table first to perform a delete of the old view row. When doing so, the memory used for the read is tracked by the user read concurrency semaphore. When we have a large number of such reads, we may use up all of the semaphore units, causing the following reads to be queued. When we have some user reads coming at the same time, these reads can have very high latency due to the write workload on the base table. We want to avoid this, so that the write workload doesn't have a high impact on the latency of the read workload. This is fixed in this patch by adding a separate read concurrency semaphore just for view update read-before-writes. With the new semaphore, even if there are many view update read-before-writes, they will be queued on a different semaphore than the user reads, and they won't impact their latency. The second issue fixed by this patch is the concurrency of the view updates that is currently unlimited. Because of that view updates may take up so much memory that they we may run out of memory. This is fixed by using the read admission on the view update concurrency semaphore. This limits the number of concurrent view update reads to max_count_concurrent_view_update_reads, all other incoming view update reads are queued using just a small chunk of memory. Without this, the reads would also get queued after exceeding view_update_reader_concurrency_semaphore_serialize_limit_multiplier, but they would take much more memory while staying in the queue. The new semaphore has half the capacity of the regular user read concurrency semahpore and is currently used only for user writes - is't used independently of the scheduling group on which we base the read semaphore selection, but we use a different code path for streaming (not database::do_apply) and we shouldn't have view updates in system writes or during compaction. Fixes https://github.com/scylladb/scylladb/issues/8873 Fixes https://github.com/scylladb/scylladb/issues/15805	2024-10-21 11:02:06 +02:00
Kefu Chai	5255f18c35	date: do not put space before literal operator when compiling date.h, clang 20 complains: ``` /home/kefu/.local/bin/clang++ -DDEBUG -DDEBUG_LSA_SANITIZER -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -isystem /home/kefu/dev/scylladb/build/rust -isystem /home/kefu/dev/scylladb/seastar/include -isystem /home/kefu/dev/scylladb/build/Debug/seastar/gen/include -isystem /usr/include/p11-kit-1 -isystem /home/kefu/dev/scylladb/abseil -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb/build=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -std=c++23 -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -DSEASTAR_API_LEVEL=7 -DSEASTAR_BUILD_SHARED_LIBS -DSEASTAR_SSTRING -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_DEBUG -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEBUG_PROMISE -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_TYPE_ERASE_MORE -DFMT_SHARED -DWITH_GZFILEOP -MD -MT lang/CMakeFiles/lang.dir/Debug/lua.cc.o -MF lang/CMakeFiles/lang.dir/Debug/lua.cc.o.d -o lang/CMakeFiles/lang.dir/Debug/lua.cc.o -c /home/kefu/dev/scylladb/lang/lua.cc In file included from /home/kefu/dev/scylladb/lang/lua.cc:18: /home/kefu/dev/scylladb/utils/date.h:836:34: error: identifier '_d' preceded by whitespace in a literal operator declaration is deprecated [-Werror,-Wdeprecated-literal-operator] 836 \| CONSTCD11 date::day operator "" _d(unsigned long long d) NOEXCEPT; \| ~~~~~~~~~~~~^~ \| operator""_d ``` because, in [CWG2521](https://wg21.link/CWG2521), it proposes that compiler should consider ```c++ string operator "" _i18n(const char*, std::size_t); // OK, deprecated ``` as "OK, deprecated". and Clang implemented this proposal, as it was accepted by C++23. since scylladb uses C++23 standard. let's remove the space between `"` and `_` to be more compliant to the C++23 standard and to silence the warning, which is taken as an error. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21194	2024-10-21 11:21:52 +03:00
Kefu Chai	8355056453	build: cmake: expose and use the path to iotune correctly in `415c83fa`, we introduced a regression which broke the build of target of "package". because - the IMPORT_LOCATION_<CONFIG> of the imported target of "Seastar::iotune" includes a literal `$<CONFIG>` - we retrieve the property named "IMPORTED_LOCATION" from this target. but value of this property is empty. so, when we copied this file, the "src" parameter passed to `cmake -E copy` is actually an empty string. in this change, we - set the `IMPORTED_LOCATION_${CONFIG}` property with a correct path. - retrieve the property with the right approach -- to use `TARGET_FILE` generator expression. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21181	2024-10-21 10:32:51 +03:00
Avi Kivity	b5a1173880	utils: small_vector: support from_range_t std::ranges::to<>() has a little protocol with containers to allow them to optimize their construction from ranges. Implement it for small_vector. It optimizes ranges that can have their size determined quickly, or that can be traversed twice to determine the size by reserving up front. Single-pass ranges (std::ranges::input_range) use the less efficient push_back method. A unit test (which fails without the new constructor) is added. Closes scylladb/scylladb#21094	2024-10-21 09:31:38 +03:00
Kefu Chai	d28d64f7fe	service: remove extraneous space in `#pragma once` to be more consistent with the rest of the tree. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21188	2024-10-20 20:27:38 +03:00
Avi Kivity	c3be2489ce	treewide: drop includes of <boost/range/adaptors.hpp> This includes way too much, including <boost/regex.hpp>, which is huge. Drop includes of adaptors.hpp and replace by what is needed. Closes scylladb/scylladb#21187	2024-10-20 17:17:11 +03:00
Aleksandra Martyniuk	29c2d4e7eb	tasks: add comments about map_each_task safety Closes scylladb/scylladb#21172	2024-10-19 21:16:38 +03:00
Avi Kivity	9a521c25b5	Merge 'test/boost: stop using ranges::to()' from Kefu Chai now that we are able to use ranges library provided by the C++ standard library. there is no need to use the homebrew `ranges::to()`. in this series, we - switch to `std::ranges::to()` in favor of `ranges::to()`. - and drop the unused `utils/ranges.hh` header file. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#21182 * github.com:scylladb/scylladb: utils: remove unused ranges.hh test/boost: stop using ranges::to()	2024-10-19 16:57:51 +03:00
Kefu Chai	c5e666b7b1	column_computation.hh: include used header when building the check-header target, we have following failure: ``` FAILED: CMakeFiles/check-headers-scylla-main.dir/Dev/check-headers/column_computation.hh.cc.o /home/kefu/.local/bin/clang++ -DDEVEL -DSCYLLA_BUILD_MODE=dev -DSCYLLA_ENABLE_ERROR_INJECTION -DSCYLLA_ENABLE_PREEMPTION_SOURCE -DSEASTAR_ENABLE_ALLOC_FAILURE_INJECTION -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Dev\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -isystem /home/kefu/dev/scylladb/seastar/include -isystem /home/kefu/dev/scylladb/build/Dev/seastar/gen/include -isystem /usr/include/p11-kit-1 -isystem /home/kefu/dev/scylladb/abseil -O2 -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb/build=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -Wno-unused-const-variable -Wno-unused-function -Wno-unused-variable -std=c++23 -Werror=unused-result -fstack-clash-protection -DSEASTAR_API_LEVEL=7 -DSEASTAR_BUILD_SHARED_LIBS -DSEASTAR_SSTRING -DSEASTAR_ENABLE_ALLOC_FAILURE_INJECTION -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_TYPE_ERASE_MORE -DFMT_SHARED -DWITH_GZFILEOP -MD -MT CMakeFiles/check-headers-scylla-main.dir/Dev/check-headers/column_computation.hh.cc.o -MF CMakeFiles/check-headers-scylla-main.dir/Dev/check-headers/column_computation.hh.cc.o.d -o CMakeFiles/check-headers-scylla-main.dir/Dev/check-headers/column_computation.hh.cc.o -c /home/kefu/dev/scylladb/build/check-headers/column_computation.hh.cc /home/kefu/dev/scylladb/build/check-headers/column_computation.hh.cc:24:37: error: no template named 'unique_ptr' in namespace 'std' 24 \| using column_computation_ptr = std::unique_ptr<column_computation>; \| ~~~~~^ /home/kefu/dev/scylladb/build/check-headers/column_computation.hh.cc:40:12: error: unknown type name 'column_computation_ptr'; did you mean 'column_computation'? 40 \| static column_computation_ptr deserialize(bytes_view raw); \| ^~~~~~~~~~~~~~~~~~~~~~ \| column_computation ``` it turns out we failed to include `<memory>`. in this change, we include `<memory>` so that this header is self-contained. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21185	2024-10-19 16:56:02 +03:00
Raphael S. Carvalho	dfc217f99a	locator: Always preserve balancing_enabled in tablet_metadata::copy() When there are zero tablets, tablet_metadata::_balancing_enabled is ignored in the copy. The property not being preserved can result in balancer not respecting user's wish to disable balancing when a replica is created later on. Fixes #21175. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Closes scylladb/scylladb#21177	2024-10-19 14:51:36 +02:00
Kefu Chai	4d4b0b35b7	utils: remove unused ranges.hh now that this header is not used, let's drop it. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-19 13:21:20 +08:00
Kefu Chai	85518463a9	test/boost: stop using ranges::to() now that we are able to use ranges library provided by the C++ standard library. there is no need to use the homebrew `ranges::to()`. in this change, we switch to `std::ranges::to()` in favor of `ranges::to()`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-19 13:21:20 +08:00
Kefu Chai	5c0db8a49e	sstable_directory: remove extraneous semicolon one semicolon is enough to mark the end of a statement. so let's remove the extraneous one. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21171	2024-10-18 21:58:04 +03:00
Kefu Chai	e2b18eb7eb	data_dictionary: compose the location with "/" in `787ea4b1`, we construct a new `storage_options` for each sstable to be restored. the `location` of the new `storage_option` instances is composed of the configured `prefix` and the dirname of each toc component. but instead of separating them with "/", we just concatenate them. this breaks the test if the specified key representing toc components includes "dirname" in them. in this change - data_directory: instead of using "{prefix}{dirname}", we use "{prefix}/{dirname}". - test/object_store: update the existing test to add a suffix in the keys of the toc objects to mimic the typical use case. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21170	2024-10-18 21:57:56 +03:00
Lakshmi Narayanan Sreethar	afad1b3c85	topology-custom: add test to verify tombstone gc in read path Co-authored-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-10-18 19:20:03 +05:30
Lakshmi Narayanan Sreethar	5a93277904	replica/table: check memtable before discarding tombstone during read On the read path, the compacting reader is applied only to the sstable reader. This can cause an expired tombstone from an sstable to be purged from the request before it has a chance to merge with deleted data in the memtable leading to data resurrection. Fix this by checking the memtables before deciding to purge tombstones from the request on the read path. A tombstone will not be purged if a key exists in any of the table's memtables with a minimum live timestamp that is lower than the maximum purgeable timestamp. Fixes #20916 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-10-18 19:19:58 +05:30
Lakshmi Narayanan Sreethar	6a357b55e3	compaction_group: track maximum timestamp across all sstables This will be used in a following patch to decide if the compacting reader has to check the memtables before purging a tombstone. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-10-18 19:19:11 +05:30
Pavel Emelyanov	b11d50f591	Merge 'multishard reader: make it safe to create with admitted permits' from Botond Dénes Passing an admitted permit -- i.e. one with count resources on it -- to the multishard reader, will possibly result in a deadlock, because the permit of the multishard reader is destroyed after the permits of its child readers. Therefore its semaphore resources won't be automatically released until children acquire their own resources. This creates a dependency (an edge in the "resource allocation graph"), where the semaphore used by the multishard reader depends on the semaphores used by children. When such dependencies create a cycle, and permits are acquired by different reads in just the right order, a deadlock will happen. Users of the multishard reader have to be aware of this gotcha -- and of course they aren't. This is small wonder, considering that not even the documentation on the multishard reader mentions this problem. To work around this, the user has to call `reader_permit::release_base_resources()` on the permit, before passing it to the multishard reader. On multiple occasions, developers (including the very author of the multishard reader), forgot or didn't know about this and this resulted in deadlocks down the line. This is a design-flaw of the multishard reader, which is addressed in this PR, after which, it is safe to pass admitted or not admitted permits to the multishard reader, it will handle the call to `release_base_resources()` if needed. After fixing the problem in the multishard reader, the existing calls to `release_base_resources()` on permits passed to multishard readers are removed. A test is added which reproduces the problem and ensures we don't regress. Refs: https://github.com/scylladb/scylladb/issues/20885 (partial fix, there is another deadlock in that issue, which this PR doesn't fix) This fixes (indirectly) a regression introduced by `d98708013c` so it has to be backported to 6.2 Closes scylladb/scylladb#21058 * github.com:scylladb/scylladb: test/boost/mutation_test: add test for multishard permit safety test/lib/reader_lifecycle_policy: add semaphore factory to constructor test/lib/reader_lifecycle_policy: rename factory_function repair/row_level: drop now unneeded release_base_resource() calls readers/multishard: make multishard reader safe to create with admitted permits	2024-10-18 13:30:21 +03:00
Pavel Emelyanov	280cd23c13	Merge 'Allow specifying TLS options with internode_encryption=none + add "transitional" mode' from Calle Wilund Fixes #18903 Adds a "transitional" internode encryption mode, under which all _outgoing_ RPC connections will use TLS, but we will still accept any incoming non-tls connection. This allows an operator to perform a move to TLS RPC without cluster downtime: 1. For each server, add certificate etc options to server_encryption_options + internode_encryption=none + set ssl_storage_port + restart (rolling) 2. For each server, set internode_encryption=transitional + RR 3. For each server, set internode_encryption=all + RR Closes scylladb/scylladb#18939 * github.com:scylladb/scylladb: test::topology: Add test for TLS upgrade and downgrade of internode encryption docs: Add internode_encryption=transitional documentation messaging_service: Add "transitional" internode encryptipn mode messaging_service: Create TLS connector even if internode_enc=none when certs set	2024-10-18 11:01:07 +03:00
Avi Kivity	1bbd1436b4	types: move from boost ranges to standard ranges Reduce depdendency load. tuple_deserializing_iterator gained a default constructor so it matches iterator constraints. Closes scylladb/scylladb#21029	2024-10-18 11:00:49 +03:00
Botond Dénes	b6da82dba3	Merge 'build: build seastar as an external project' from Kefu Chai before this change, scylla's CMake-based system consumes Seastar library by including it directly. but this failed to address the needs of linking against Seastar shared libraries in Debug and Dev builds, while linking against the static libraries in other builds. because Seastar uses `BUILD_SHARED_LIBS` CMake variable to determine if it builds shared libraries. and we cannot assign different values to this CMake variable based on current configure type -- CMake does not support. see https://gitlab.kitware.com/cmake/cmake/-/issues/19467 in order to address this problem, we have a couple possible solutions: - to enable Seastar to build both shared and static libraries in a pass. without sacrificing the performance, we have to build all object files twice: once with -fPIC, once without. in order to accompolish this goal, we need to develop a machinary to populate the same settings to these two builds. this would complicate the design of Seastar's building system further. - to build Seastar libraries twice in scylla, we could use the ExternalProject module to implement this. but it'd be complicate to extract the compile options, and link options previously populated by Seastar's targets with CMake -- we would have to replicate all of them in scylla. this is out of the question. - to build Seastar libraries twice before building scylla, and let scylla to consume them using CMake config files or .pc files. this is a compromise. it enables scylla to drive the build of Seastar libraries and to consume the compile options and link options. the downside is: * the generated compilation database (compile_commands.json) does not include the commands building Seastar anymore. * the building system of scylla does not have finer graind control on the building process of seastar. for instance, we cannot specify the build dependency to a certain seastar library, and just build it instead of building the whole seastar project. turns out the last approach is the best one we can have at this moment. this is also the approach used by the existing `configure.py`. in this change, we - add FindSeastar.cmake to * detect the preconfigured Seastar builds, and * extract the build options from .pc files * expose library targets to be consumed by parent project - add Seastar as an external project, so we can build it from the parent project. this is atypical compared to standard ExternalProject usage: - Seastar's build system should already be configured at this point. - We maintain separate project variants for each configuration type. Benefits of this approach: - Allows the parent project to consume the compile options exposed by .pc file. as the compile options vary from one config to another. - Allows application of config-specific settings - Enables building Seastar within the parent project's build system - Facilitates linking of artifacts with the external project target, establishing proper dependencies between them we will update `configure.py` to merge the compilation database of scylla and seastar. Refs scylladb/scylladb#2717 --- this is a CMake-related change, hence no need to backport. Closes scylladb/scylladb#21131 * github.com:scylladb/scylladb: build: cmake: use GENERATOR_IS_MULTI_CONFIG property to detect mult-config build: cmake: consume Seastar using its .pc files build: do not use `mode` as the index into `modes` build: cmake: detect and link against GnuTLS library build: cmake: detect and link against yaml-cpp build: cmake: link Seastar with Seastar::<COMPONENT> build: cmake: define CMake generate helper funcs in scylla	2024-10-18 09:42:59 +03:00
Amnon Heiman	09fa625672	scripts/get_description.py: param_mapping was missing get_description.py was moved from a standalone script to a library. During the transition, param_mapping was not included in the script option. This patch makes it possible to use the file as a standalone script again.	2024-10-18 08:58:04 +03:00
Amnon Heiman	10af854ec4	scripts/metrics-config.yml: no need to get metrics from the tests	2024-10-18 08:57:53 +03:00
Kefu Chai	b5f5a963ca	build: do not pass Seastar_CXX_DIALECT=gnu++23 when building Seastar Seastar now respect CMAKE_CXX_STANDARD in favor of Seastar_CXX_DIALECT, which has been dropped in Seastar's commit of 60bc8603bd438232614e9b3dcd7537dc83c85206 . Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21130	2024-10-18 08:57:23 +03:00
Botond Dénes	6811411288	Merge 'Sanitize commitlog API endpoints' from Pavel Emelyanov Endpoints are registered next to the service they use, and the unregistration deferred action is created right after it. When registered, the service in question is passed as argument and then captured by enpoints lambdas. This makes sure that service is not used by endpoints after being stopped. That's not so for commitlog endpoints. These are registered in several places, and /commitlog "function" is not unregistered on stop. This patch fixes some of this misbehavior, in particular: - adds unregistration of commitlog API function - uses sharded<database>& argument in endpoints instead of ctx.db - moves some endpoints from storage_service.cc to commitlog.cc Closes scylladb/scylladb#21053 * github.com:scylladb/scylladb: api: Use captured database, not the one from ctx api: Pass sharded<database> to commitlog endpoints registration api: Move commitlog-related from storage_service.cc api: Unset commitlog API endpoints api: Extract set_server_commitlog() from set_server_done()	2024-10-18 08:56:13 +03:00
Botond Dénes	568b767ec3	Merge 'schema: convert from boost ranges to std ranges' from Avi Kivity To reduce dependency load, change uses of boost ranges to std::ranges. The first patch is preparation, replacing a construct that isn't easy to support with std ranges with something simpler. No backport as this is a code cleanup. Closes scylladb/scylladb#21122 * github.com:scylladb/scylladb: schema: replace boost ranges with std ranges schema: precompute all_columns_in_select_order()	2024-10-18 08:42:50 +03:00
Pavel Emelyanov	df6991edd3	test: Do not duplicate sstable twice The statistics_rewrite test case copies an sstable from resources two times: - first time -- explicitly by listing resource components and copying files to the test temp dir - second time -- implicitly, by calling create_links() linking copied files by new set in the staging/ subdirectory The 2nd step is not needed and the history of changes justifies that. The test itself appeared with `70b793e4d3` and it only contained the 2nd "copying" -- test linked files from resource directory and then worked in the newly created set. Later, commit `59c57861ae` added the first step and copied the files from resource into test temp dir. At this point linking copied files because pointless, but was preserved. Let's remove it now. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21097	2024-10-18 08:31:08 +03:00
Kefu Chai	26a5a00b20	interval: include used header when building the tree with Clang-20 and libstdc++ shippped with GCC-14.2, we have following build failure: ``` /home/kefu/dev/scylladb/interval.hh:638:14: error: no member named 'sort' in namespace 'std' 638 \| std::sort(intervals.begin(), intervals.end(), [&](auto&& r1, auto&& r2) { \| ~~~~~^ /home/kefu/dev/scylladb/interval.hh:691:21: error: no member named 'upper_bound' in namespace 'std' 691 \| return std::upper_bound(r.begin(), r.end(), value, std::forward<LessComparator>(cmp)); \| ~~~~~^ /home/kefu/dev/scylladb/interval.hh:723:18: error: no member named 'minmax' in namespace 'std'; did you mean 'fminmag'? 723 \| auto p = std::minmax(_interval, other._interval, [&cmp] (auto&& a, auto&& b) { \| ^~~~~~~~~~~ \| fminmag ``` it turns out we failed to include the used header. in this change, we include `<algorithm>` so that this header is self-contained. after this change, the build passes. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21168	2024-10-18 08:26:27 +03:00
Kefu Chai	e73b0c942f	build: cmake: use GENERATOR_IS_MULTI_CONFIG property to detect mult-config this is more reliable way to check if we are configured to use a mult-config generator. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-18 08:36:52 +08:00
Kefu Chai	415c83fa67	build: cmake: consume Seastar using its .pc files before this change, scylla's CMake-based system consumes Seastar library by including it directly. but this failed to address the needs of linking against Seastar shared libraries in Debug and Dev builds, while linking against the static libraries in other builds. because Seastar uses `BUILD_SHARED_LIBS` CMake variable to determine if it builds shared libraries. and we cannot assign different values to this CMake variable based on current configure type -- CMake does not support. see https://gitlab.kitware.com/cmake/cmake/-/issues/19467 in order to address this problem, we have a couple possible solutions: - to enable Seastar to build both shared and static libraries in a pass. without sacrificing the performance, we have to build all object files twice: once with -fPIC, once without. in order to accompolish this goal, we need to develop a machinary to populate the same settings to these two builds. this would complicate the design of Seastar's building system further. - to build Seastar libraries twice in scylla, we could use the ExternalProject module to implement this. but it'd be complicate to extract the compile options, and link options previously populated by Seastar's targets with CMake -- we would have to replicate all of them in scylla. this is out of the question. - to build Seastar libraries twice before building scylla, and let scylla to consume them using CMake config files or .pc files. this is a compromise. it enables scylla to drive the build of Seastar libraries and to consume the compile options and link options. the downside is: * the generated compilation database (compile_commands.json) does not include the commands building Seastar anymore. * the building system of scylla does not have finer graind control on the building process of seastar. for instance, we cannot specify the build dependency to a certain seastar library, and just build it instead of building the whole seastar project. turns out the last approach is the best one we can have at this moment. this is also the approach used by the existing `configure.py`. in this change, we - add FindSeastar.cmake to * detect the preconfigured Seastar builds, and * extract the build options from .pc files * expose library targets to be consumed by parent project - add Seastar as an external project, so we can build it from the parent project. BUILD_AWAYS is set to ensure that Seastar is rebuilt, as scylla developers are expected to modify Seastar occasionally. since the change in Seastar's SOURCE_DIR is not detectable via the ExternalProject, we have to rebuild it. this is atypical compared to standard ExternalProject usage: - Seastar's build system should already be configured at this point. - We maintain separate project variants for each configuration type. Benefits of this approach: - Allows the parent project to consume the compile options exposed by .pc file. as the compile options vary from one config to another. - Allows application of config-specific settings - Enables building Seastar within the parent project's build system - Facilitates linking of artifacts with the external project target, establishing proper dependencies between them - preserve the existing machinery of including Seastar only when building without multi-config generator. this allows users who don't use mult-config generator to build Seastar in-the-tree. the typical use case is the CI workflows performing the static analysis. we will update `configure.py` to merge the compilation database of scylla and seastar. Refs scylladb/scylladb#2717 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-18 08:36:52 +08:00
Kefu Chai	7cb74df323	build: do not use `mode` as the index into `modes` before this change, in `configure_seastar()`, we use `mode` as a component in the build directory, and use it as the index into `modes` dict. but in a succeeding commit, we will reuse `configure_seastar()` when preparing for the CMake-based building system, in which, `mode` will be the CMake configure type, like "Debug" instead of scylla's build mode, like "debug". to be prepared for this change, let's use `mode_config` directly. it's identical to `modes[mode]`. this also improves the readability. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-18 08:36:52 +08:00
Kefu Chai	1bd2ed7826	build: cmake: detect and link against GnuTLS library before this change, in the CMake-based building system, we rely on Seastar to provide this linkage, but this is wrong and fragile. as Seastar is not supposed to expose and provide GnuTLS symbols. that's why we have following build failure: ``` : && /home/kefu/.local/bin/clang++ -g -Og -g -gz -Xlinker --build-id=sha1 --ld-path=ld.lld -dynamic-linker=/////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////////lib64/ld-linux-x86-64.so.2 /home/kefu/dev/scylladb/build/Debug/seastar/libseastar.so -fsanitize=address -fsanitize=undefined /usr/lib64/libboost_program_options.so /usr/lib64/libboost_thread.so /usr/lib64/libcares.so /usr/lib64/libfmt.so.11.0.2 -L/usr/lib64 -llz4 CMakeFiles/scylla_version.dir/Debug/release.cc.o CMakeFiles/scylla.dir/Debug/main.cc.o -o Debug/scylla -L/home/kefu/dev/scylladb/idl/absl::headers -Wl,-rpath,/home/kefu/dev/scylladb/idl/absl::headers:/home/kefu/dev/scylladb/build/Debug/seastar Debug/libscylla-main.a api/Debug/libapi.a alternator/Debug/libalternator.a db/Debug/libdb.a cdc/Debug/libcdc.a compaction/Debug/libcompaction.a cql3/Debug/libcql3.a data_dictionary/Debug/libdata_dictionary.a gms/Debug/libgms.a index/Debug/libindex.a lang/Debug/liblang.a message/Debug/libmessage.a mutation/Debug/libmutation.a mutation_writer/Debug/libmutation_writer.a raft/Debug/libraft.a readers/Debug/libreaders.a redis/Debug/libredis.a repair/Debug/librepair.a replica/Debug/libreplica.a schema/Debug/libschema.a service/Debug/libservice.a sstables/Debug/libsstables.a streaming/Debug/libstreaming.a test/perf/Debug/libtest-perf.a tools/Debug/libtools.a transport/Debug/libtransport.a types/Debug/libtypes.a utils/Debug/libutils.a Debug/seastar/libseastar.so /usr/lib64/libyaml-cpp.so /usr/lib64/libboost_program_options.so.1.83.0 test/lib/Debug/libtest-lib.a -Xlinker --push-state -Xlinker --whole-archive auth/Debug/libscylla_auth.a -Xlinker --pop-state /usr/lib64/libcrypt.so cdc/Debug/libcdc.a compaction/Debug/libcompaction.a mutation_writer/Debug/libmutation_writer.a -Xlinker --push-state -Xlinker --whole-archive dht/Debug/libscylla_dht.a -Xlinker --pop-state index/Debug/libindex.a -Xlinker --push-state -Xlinker --whole-archive locator/Debug/libscylla_locator.a -Xlinker --pop-state message/Debug/libmessage.a gms/Debug/libgms.a sstables/Debug/libsstables.a readers/Debug/libreaders.a schema/Debug/libschema.a -Xlinker --push-state -Xlinker --whole-archive tracing/Debug/libscylla_tracing.a -Xlinker --pop-state Debug/libscylla-main.a -Xlinker --push-state -Xlinker --whole-archive Debug/libscylla-zstd.a -Xlinker --pop-state /usr/lib64/libzstd.so abseil/absl/strings/Debug/libabsl_cord.a abseil/absl/strings/Debug/libabsl_cordz_info.a abseil/absl/strings/Debug/libabsl_cord_internal.a abseil/absl/strings/Debug/libabsl_cordz_functions.a abseil/absl/strings/Debug/libabsl_cordz_handle.a abseil/absl/crc/Debug/libabsl_crc_cord_state.a abseil/absl/crc/Debug/libabsl_crc32c.a abseil/absl/crc/Debug/libabsl_crc_internal.a abseil/absl/crc/Debug/libabsl_crc_cpu_detect.a abseil/absl/strings/Debug/libabsl_str_format_internal.a service/Debug/libservice.a node_ops/Debug/libnode_ops.a service/Debug/libservice.a node_ops/Debug/libnode_ops.a raft/Debug/libraft.a repair/Debug/librepair.a streaming/Debug/libstreaming.a replica/Debug/libreplica.a abseil/absl/container/Debug/libabsl_raw_hash_set.a abseil/absl/hash/Debug/libabsl_hash.a abseil/absl/hash/Debug/libabsl_city.a abseil/absl/types/Debug/libabsl_bad_variant_access.a abseil/absl/hash/Debug/libabsl_low_level_hash.a abseil/absl/types/Debug/libabsl_bad_optional_access.a abseil/absl/container/Debug/libabsl_hashtablez_sampler.a abseil/absl/profiling/Debug/libabsl_exponential_biased.a abseil/absl/synchronization/Debug/libabsl_synchronization.a abseil/absl/debugging/Debug/libabsl_stacktrace.a abseil/absl/synchronization/Debug/libabsl_graphcycles_internal.a abseil/absl/synchronization/Debug/libabsl_kernel_timeout_internal.a abseil/absl/debugging/Debug/libabsl_symbolize.a abseil/absl/debugging/Debug/libabsl_debugging_internal.a abseil/absl/base/Debug/libabsl_malloc_internal.a abseil/absl/debugging/Debug/libabsl_demangle_internal.a abseil/absl/time/Debug/libabsl_time.a abseil/absl/strings/Debug/libabsl_strings.a abseil/absl/strings/Debug/libabsl_strings_internal.a abseil/absl/strings/Debug/libabsl_string_view.a abseil/absl/base/Debug/libabsl_throw_delegate.a abseil/absl/numeric/Debug/libabsl_int128.a abseil/absl/base/Debug/libabsl_base.a abseil/absl/base/Debug/libabsl_raw_logging_internal.a abseil/absl/base/Debug/libabsl_log_severity.a abseil/absl/base/Debug/libabsl_spinlock_wait.a -lrt abseil/absl/time/Debug/libabsl_civil_time.a abseil/absl/time/Debug/libabsl_time_zone.a -lsystemd /usr/lib64/libz.so /usr/lib64/libdeflate.so types/Debug/libtypes.a utils/Debug/libutils.a /usr/lib64/libyaml-cpp.so /usr/lib64/libcryptopp.so /usr/lib64/libboost_regex.so.1.83.0 /usr/lib64/libicui18n.so /usr/lib64/libicuuc.so -ldl /usr/lib64/libboost_unit_test_framework.so.1.83.0 Debug/seastar/libseastar_perf_testing.so /usr/lib64/libjsoncpp.so.1.9.5 db/Debug/libdb.a data_dictionary/Debug/libdata_dictionary.a cql3/Debug/libcql3.a transport/Debug/libtransport.a cql3/Debug/libcql3.a transport/Debug/libtransport.a lang/Debug/liblang.a /usr/lib64/liblua-5.4.so -lm rust/Debug/libwasmtime_bindings.a rust/librust_combined.a /usr/lib64/libsnappy.so.1.2.1 mutation/Debug/libmutation.a Debug/seastar/libseastar.so /usr/lib64/liblz4.so /usr/lib64/libxxhash.so && : ld.lld: error: undefined symbol: gnutls_hmac_fast >>> referenced by aws_sigv4.cc:21 (/home/kefu/dev/scylladb/utils/aws_sigv4.cc:21) >>> aws_sigv4.cc.o:(utils::aws::hmac_sha256(std::basic_string_view<char, std::char_traits<char>>, std::basic_string_view<char, std::char_traits<char>>)) in archive utils/Debug/libutils.a ld.lld: error: undefined symbol: gnutls_strerror >>> referenced by aws_sigv4.cc:23 (/home/kefu/dev/scylladb/utils/aws_sigv4.cc:23) >>> aws_sigv4.cc.o:(utils::aws::hmac_sha256(std::basic_string_view<char, std::char_traits<char>>, std::basic_string_view<char, std::char_traits<char>>)) in archive utils/Debug/libutils.a ``` in this change, we detect this library, and link its caller against it. this addresses the link failure. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-18 08:36:52 +08:00
Kefu Chai	fc8212483e	build: cmake: detect and link against yaml-cpp in main.cc, we use yaml-cpp library directly. so we are obliged to detect this library in scylla and link against it instead of relying on other library to do this. currently, Seastar detects it and pulls in yaml-cpp for us, but we should not take this for granted and rely on this. in this change, we detect and link against yaml-cpp to make this dependency explicit. the same applies to the "utils" library. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-18 08:36:52 +08:00
Kefu Chai	2e4be56112	build: cmake: link Seastar with Seastar::<COMPONENT> before this change, we link against the targets defined in Seastar's source tree. but these targets are not part of Seastar's public interface -- they are not exposed by Seastar's CMake config files. so, let link against the target names qualified by the library module name. this also prepares for the transition to using Seastar without including it directly. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-18 08:36:52 +08:00
Kefu Chai	b2dc261841	build: cmake: define CMake generate helper funcs in scylla before this change, we assume that scylla's CMake script includes Seastar's CMake script. but we are going to consume Seastar using its .pc files or its CMake config files instead of including it directly. more over these helper functions are not part of Seastar's public interface. actually the same applies to the `check_headers()` helper, which was adapted from seastar's CheckHeaders.cmake. so to be prepared for this change, let's define these generate helper functions in scylla. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-18 08:36:52 +08:00
Avi Kivity	f4acaa5473	cql3: index_target: forward declare boost::regex No need to burden everyone with the full boost::regex code. Closes scylladb/scylladb#21148	2024-10-17 19:14:40 +02:00
Botond Dénes	e1d8cddd09	test/boost/mutation_test: add test for multishard permit safety Add a test checking that the multishard reader will not deadlock, when created with an admitted permit, on a semaphore with a single count resource.	2024-10-17 08:47:50 -04:00
Botond Dénes	5a3fd69374	test/lib/reader_lifecycle_policy: add semaphore factory to constructor Allowing callers to specify how the semaphore is created and stopped, instead of doing so via boolean flags like it is done currently. This method doesn't scale, so use a factory instead.	2024-10-17 08:47:50 -04:00
Botond Dénes	c8598e21e8	test/lib/reader_lifecycle_policy: rename factory_function To reader_factor_function. We are about to add a new factory function parameters, so the current factory_function has to be renamed to something more specific.	2024-10-17 08:47:50 -04:00
Botond Dénes	76a5ba2342	repair/row_level: drop now unneeded release_base_resource() calls The multishard reader now does this itself, no need to do it here.	2024-10-17 08:47:50 -04:00
Botond Dénes	218ea449a5	readers/multishard: make multishard reader safe to create with admitted permits Passing an admitted permit -- i.e. one with count resources on it -- to the multishard reader, will possibly result in a deadlock, because the permit of the multishard reader is destroyed after the permits of its child readers. Therefore its semaphore resources won't be automatically released until children acquire their own resources. This creates a dependency (an edge in the "resource allocation graph"), where the semaphore used by the multishard reader depends on the semaphores used by children. When such dependencies create a cycle, and permits are acquired by different reads in just the right order, a deadlock will happen. Users of the multishard reader have to be aware of this gotcha -- and of course they aren't. This is small wonder, considering that not even the documentation on the multishard reader mentions this problem. To work around this, the user has to call `reader_permit::release_base_resources()` on the permit, before passing it to the multishard reader. On multiple occasions, developers (including the very author of the multishard reader), forgot or didn't know about this and this resulted in deadlocks down the line. This is a design-flaw of the multishard reader, which is addressed in this patch, after which, it is safe to pass admitted or not admitted permits to the multishard reader, it will handle the call to `release_base_resources()` if needed.	2024-10-17 08:45:21 -04:00
Raphael S. Carvalho	f3ab5e1f1e	tests: Fix perf test for load balancer Broken after introduction of zero-token nodes. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Closes scylladb/scylladb#21156	2024-10-17 14:02:31 +02:00
Kamil Braun	f02afefd34	Merge 'raft: consider the gossiper state then sending the group0 state id' from Emil Maskovsky Skip the advertisement of the group0 state id in case the gossiper is not active (ready). Sending the application state when the gossiper is not active caused a warning being shown in the log about the local endpoint not being found in the gossiper endpoint state map on a (graceful) node restart. The local endpoint is initialized on the gossiper startup, so we skip the state id advertisement until the startup is finished. Fixes: scylladb/scylladb#21117 No backport: Fixes an issue that is currently only present in master Closes scylladb/scylladb#21119 * github.com:scylladb/scylladb: raft: consider the gossiper state then sending the group0 state id raft: add the test for GROUP0_STATE_ID gossip application state	2024-10-17 13:41:15 +03:00
Kefu Chai	5ef0cbb693	tools/scylla-nodetool: s/vm.count()/vm.contains()/ this change is created in the same spirit of `0104c7d3`, which used `std::map::contains()` in the place of `std::map::count()` when checking for the existence of a paramter with given name for better readability. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21158	2024-10-17 13:41:15 +03:00
Alexey Novikov	b965729f0a	replica: implement memtable_flush_period_in_ms schema option implement cassandra original schema option memtable_flush_period_in_ms: Milliseconds before memtables associated with the table are flushed. there are few things concerning this patch: * milliseconds look strange and scary for this option. Unlike Cassandra we use 60000ms (1min) minimum value for this option. * This is limitation of Cassandra but it is impossible to set this option for system tables. However sometimes it could be very useful to use automatic flushing for such a tables: some system tables have small traffic and as a result prevent tombstone garbage collection. Fixes #20270 Closes scylladb/scylladb#20999	2024-10-17 13:41:15 +03:00
Anna Stuchlik	b54ce3b0c0	doc: remove the redundant raw:: html directive This commit removes the raw:: html directive (with the exception of an embedded animation) because: - It is not supported by the dark theme and looks bad. - It's a legacy directive, and we no longer need it on index pages. Fixes https://github.com/scylladb/scylladb/issues/20881 Closes scylladb/scylladb#21062	2024-10-17 13:41:15 +03:00
Kefu Chai	d7f315ef63	tool/scylla-nodetool: check for positional argument passed to "restore" before this change, if no positional arguments are passed to "restore" subcommand, the tool fails with following error message: ``` error running operation: boost::wrapexcept<boost::bad_any_cast> (boost::bad_any_cast: failed conversion using boost::any_cast) ``` this is difficult to digest. after this change, if no sstables are specified: ``` error processing arguments: missing required parameter: sstables ``` this is slightly better from user experience's perspective. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21136	2024-10-17 13:41:15 +03:00
Kefu Chai	9355a32b5c	utils/loading_cache: s/typeof/decltype/ `typeof` is a GNU extension, and is part of C23, but it is not included by C++23. if we compile the tree with c++23 instead of gnu++23, the compilation fails like: ``` FAILED: repair/CMakeFiles/repair.dir/RelWithDebInfo/repair.cc.o /home/kefu/.local/bin/clang++ -DSCYLLA_BUILD_MODE=release -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"RelWithDebInfo\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -isystem /home/kefu/dev/scylladb/abseil -isystem /home/kefu/dev/scylladb/seastar/include -isystem /home/kefu/dev/scylladb/build/RelWithDebInfo/seastar/gen/include -isystem /usr/include/p11-kit-1 -ffunction-sections -fdata-sections -O3 -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb/build=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -mllvm -inline-threshold=2500 -fno-slp-vectorize -std=c++23 -Werror=unused-result -DSEASTAR_API_LEVEL=7 -DSEASTAR_SSTRING -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_LOGGER_TYPE_STDOUT -DFMT_SHARED -DWITH_GZFILEOP -MD -MT repair/CMakeFiles/repair.dir/RelWithDebInfo/repair.cc.o -MF repair/CMakeFiles/repair.dir/RelWithDebInfo/repair.cc.o.d -o repair/CMakeFiles/repair.dir/RelWithDebInfo/repair.cc.o -c /home/kefu/dev/scylladb/repair/repair.cc In file included from /home/kefu/dev/scylladb/repair/repair.cc:21: In file included from /home/kefu/dev/scylladb/service/storage_service.hh:19: In file included from /home/kefu/dev/scylladb/service/qos/service_level_controller.hh:19: In file included from /home/kefu/dev/scylladb/auth/service.hh:23: In file included from /home/kefu/dev/scylladb/auth/permissions_cache.hh:22: /home/kefu/dev/scylladb/utils/loading_cache.hh:754:66: error: use of undeclared identifier 'typeof'; did you mean 'typeid'? 754 \| static_assert(SectionHitThreshold <= std::numeric_limits<typeof(_touch_count)>::max() / 2, "SectionHitThreshold value is too big"); \| ^ /home/kefu/dev/scylladb/utils/loading_cache.hh:754:66: error: template argument for template type parameter must be a type 754 \| static_assert(SectionHitThreshold <= std::numeric_limits<typeof(_touch_count)>::max() / 2, "SectionHitThreshold value is too big"); \| ^~~~~~~~~~~~~~~~~~~~ /usr/lib/gcc/x86_64-redhat-linux/14/../../../../include/c++/14/limits:311:21: note: template parameter is declared here 311 \| template<typename _Tp> \| ^ 2 errors generated. ``` in this change, we trade `typeof` for a more standard compliant `decltype`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21116	2024-10-17 13:41:15 +03:00
Pavel Emelyanov	df83fe2dae	Merge 'interval: replace boost ranges with std ranges' from Avi Kivity To reduce dependency load, replace use of boost ranges with std ranges. Since std ranges are more particular about what iterators they accept, a custom iterator in size_estimates_virtual_reader has to be fixed first. No backport; code cleanup. Closes scylladb/scylladb#21143 * github.com:scylladb/scylladb: interval: change boost ranges to std ranges size_estimates_virtual_reader: make virtual_row_iterator more conforming	2024-10-17 13:41:15 +03:00
Avi Kivity	6fd219d982	sstables: generation_type: deinline from_string() This is not performance sensitive and penalizes everyone by including boost/regex.hpp. Fix by deinlining. Closes scylladb/scylladb#21147	2024-10-17 13:41:15 +03:00
Emil Maskovsky	e082fef32c	raft: remove the group0 state id handler stop check The stop assertion check in the group0 state id handler was triggering under some circumstances (stopping server during restart). In that case it might be that the stop is initiated before the server is fully initialized, and then the handler destructor is being called without calling to the `stop()` method first. This is a valid scenario. The whole `stop()` in the group0 state id handler is not necessary, as the only operation being done is cancelling the timer which is done by the timer destructor automatically anyway. There is the concern of a currently running timer callback, but it doesn't preempt (not async) so the timer shouldn't be destroyed before the callback finishes. Fixes: scylladb/scylladb#21074 Closes scylladb/scylladb#21127	2024-10-17 13:41:15 +03:00
Emil Maskovsky	3f1af268c2	raft: consider the gossiper state then sending the group0 state id Skip the advertisement of the group0 state id in case the gossiper is not active (ready). Sending the application state when the gossiper is not active caused a warning being shown in the log about the local endpoint not being found in the gossiper endpoint state map on a (graceful) node restart. The local endpoint is initialized on the gossiper startup, so we skip the state id advertisement until the startup is finished. Fixes: scylladb/scylladb#21117	2024-10-16 19:26:25 +02:00
Emil Maskovsky	65d3d4fd93	raft: add the test for GROUP0_STATE_ID gossip application state Test that the GROUP0_STATE_ID gossip application state is not causing the "endpoint_state_map does not contain endpoint" error. Refs: scylladb/scylladb#21117	2024-10-16 19:21:14 +02:00
Calle Wilund	f2ef75c3da	commitlog_test: Up timeout for large entry tests Fixes #21150 Apparently, on some CI, in debug, these tests can time out (large alloc) without actually failing what they do. Up the timeout (could consider removing as well, but...) so they hopefully pass. Closes scylladb/scylladb#21151	2024-10-16 18:13:04 +03:00
Avi Kivity	f799234c82	Update tools/java submodule (deprecation notice) * tools/java b2d025fd6b...807e991de7 (1): > README.md: add deprecation notice for java tools	2024-10-16 17:09:48 +03:00
Avi Kivity	b73f0197a8	Merge 'micro-updates to documentation development, on python-poetry' from Laszlo Ersek - `docs/Makefile`: work around python-poetry issue https://github.com/python-poetry/poetry/issues/8761 - `docs/README.md`: fix minimum poetry version No backporting needed (docs development). Closes scylladb/scylladb#21118 * github.com:scylladb/scylladb: docs/README.md: fix minimum poetry version docs/Makefile: work around python-poetry issue #8761	2024-10-16 14:16:29 +03:00
Nadav Har'El	ee0e7a7adf	mv: test that operations that should not be allowed on a view, aren't This patch adds test/cql-pytest tests which verify that all CQL operations that shouldn't be allowed on a materialized view, actually aren't: * All operations writing to a table - INSERT, UPDATE, BATCH, DELETE, and TRUNCATE - should be rejected when asked to operate on a view. * All operations with "TABLE" in their name (DROP TABLE, ALTER TABLE, DESC TABLE) should be rejected on a view - the ".. MATERIALIZED VIEW" operation should be used instead. * A materialized view cannot get materialized views or indexes of its own. All tests pass on Cassandra (Cassandra 4 or above is needed for the "DESC" test), and all but one pass on Scylla - Scylla does allow "DESC TABLE" on a materialized view, unlike Cassandra. I opened an issue to track that difference: Refs #21026 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#21028	2024-10-16 13:43:36 +03:00
Avi Kivity	d58cd262ca	interval: change boost ranges to std ranges Reduce dependency load. size_estimates_virtual_reader is adjusted due to poor boost ranges and std ranges interoperability.	2024-10-16 13:21:43 +03:00
Avi Kivity	3a75efd6d4	size_estimates_virtual_reader: make virtual_row_iterator more conforming To work with std::ranges, an iterator has to have a default constructor, and be assignable. Add the default constructor and convert references to pointers to support this.	2024-10-16 13:21:25 +03:00
Pavel Emelyanov	4a8ab9b3bc	s3/client: Restore indentation after previous patch Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-16 12:27:29 +03:00
Pavel Emelyanov	a15dfe0154	s3/client: Catch do_upload_file::upload_part() exceptions This method spawns part uploading in the background, but still may throw, e.g. preparing http request or claiming memory. In this case any outstanding part upload fibers are not waited on, and the whole do_upload_file object can be freed from under their feet. Also, the multipart upload is not aborted, thus losing track of it until g.c. happens. To fix it, catch any exception from upload_part() too, and if it happens, do what the regular upload_sink would do -- close the gate thus picking up any outstanding activity that may happen there and abort the multipart upload. Indentation is deliberately left broken Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-16 12:23:31 +03:00
Nadav Har'El	210d53070e	docs/alternator: explain service discovery HTTP requests In this patch we add to docs/new-apis.md (Alternator-specific API) a description of the service discovery HTTP requests - `/` and `/localnodes` that was previously not documented except in a design document that is unfortunately no longer available publically. The description also includes the recently added `dc` and `rack` parameters for the `/localnodes` request. Fixes #20989 Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-10-16 10:15:04 +03:00
Nadav Har'El	367e18ed4a	docs/alternator: split Alternator-specific APIs from alternator.md Before this patch, the documentation of Alternator-specific APIs (APIs which are unique to Alternator and don't exist in DynamoDB) appear as a section of the main document alternator.md. In the next patch we want to describe yet another Alternator feature and make this section even longer. But there is growing sentiment that the Alternator documentation should be split into more, shorter, pages (Refs #19822) so this patch splits the Alternator-specific API documentation into a new file, new-apis.md. There is no new content in the patch - just movement of existing content plus a reference to the new page. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-10-16 10:14:31 +03:00
Kefu Chai	32f508d450	raft: fix typo in logging message s/miminum/minimum/ Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21073	2024-10-16 06:33:43 +03:00
Avi Kivity	d59038fa93	storage_proxy: convert boost range algorithms to std::ranges Standardize on a single range library. The changes are mostly mechanical. The only exception is boost::join, which has no analog in std::ranges (rightly so, since it cannot be implemented efficiently). A variety of tricks were used to convert it: - use std::ranges::join() on an std::array of std::span (when the inputs were all contiguous) - copy to a utils::small_vector (when it is expected that there will be no allocation) - use a small_vector of pointers and iterate+dereference that Closes scylladb/scylladb#21082	2024-10-15 16:52:27 +02:00
Avi Kivity	820509026f	schema: replace boost ranges with std ranges To reduce dependency load, use std ranges instead of boost ranges. The std::ranges::{lower,upper}_bound don't support heterogeneous lookup, but a more natural solution is to use a projection to search for the name, so we use that and the custom comparator is removed. Many callers are converted as well due to poor interoperability between boost ranges and std ranges.	2024-10-15 16:42:54 +03:00
Piotr Dulikowski	a380a2efd9	test/test_view_build_status: properly wait for v2 in migration test The test_view_build_status_migration_to_v2 test case creates a new view (vt2) after peforming the view_build_status -> view_build_status_v2 migration and waits until it is built by `wait_for_view_v2` function. It works by waiting until a SELECT from view_build_status_v2 will return the expected number of rows for a given view. However, if the host parameter is unspecified, it will query only one node on each attempt. Because `view_build_status_v2` is managed via raft, queries always return data from the queried node only. It might happen that `wait_for_view_v2` fetches expected results from one node while a different node might be lagging behind the group0 coordinator and might not have all data yet. In case of test_view_build_status_migration_to_v2 this is a problem - it first uses `wait_for_view_v2` to wait for view, later it queries `view_build_status_v2` on a random node and asserts its state - and might fail because that node didn't have the newest state yet. Fix the issue by issuing `wait_for_view_v2` in parallel for all nodes in the cluster and waiting until all nodes have the most recent state. Fixes: scylladb/scylladb#21060 Closes scylladb/scylladb#21091	2024-10-15 14:57:47 +03:00
Pavel Emelyanov	63725b10a8	Merge 'cql: create default superuser if it doesn't exist' from Paweł Zakrzewski This change reorganizes the way standard_role_manager startup is handled: role_manager::ensure_superuser_is_created() is added, which returns a future that resolves once the superuser is available. We wait for this future before starting the CQL server. There is a change in behavior auth::do_after_system_ready is potentially an infinite loop, and we await its result. Fixes #10481 Reason for no backports: it's not a regresson and it's an issue that may only affect a tiny time window during the cluster startup. Closes scylladb/scylladb#20137 * github.com:scylladb/scylladb: test: test_restart_cluster: create the test auth: standard_role_manager allows awaiting superuser creation auth: coroutinize the standard_role_manager start() function auth: don't start server until the superuser is created	2024-10-15 14:56:04 +03:00
Avi Kivity	a5c37a110f	schema: precompute all_columns_in_select_order() all_columns_in_select_order() returns a complicated boost range type that has no analog in std::ranges. To ease the transition to std::ranges, precompute most of the work done in that function, and only convert pointers to references in the function itself. Since boost ranges and std::ranges don't fully interoperate, one of the user has to be adjusted.	2024-10-15 14:04:12 +03:00
Pavel Emelyanov	6e0899c2b4	data_dictionary: Replace boost ranges with std ranges Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21105	2024-10-15 13:22:08 +03:00
Laszlo Ersek	a0ffbd5bcf	docs/README.md: fix minimum poetry version Commit `2a3012db7f` ("docs/README.md: expand prerequisites list", 2022-08-31) referenced poetry release 1.12, which does not exist even today (as of this writing, the latest release is 1.8.4). The intent was probably 1.1.12. Copy the minimum version from "sphinx-scylladb-theme": 1.8.1 (see "docs/source/getting-started/installation.rst" and "docs/source/getting-started/quickstart.rst" at commit f7c26b422572). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-10-15 12:19:21 +02:00
Laszlo Ersek	e5c2d4bd1d	docs/Makefile: work around python-poetry issue #8761 Python-poetry is affected by bug <https://github.com/python-poetry/poetry/issues/8761>. Namely, if you have "keyring" <https://pypi.org/project/keyring/> installed, poetry will try to gain access to the Default collection in the (ex. GNOME) keyring, even if poetry only needs read-only access to package repositories, and even if those repos are public. Consequently, you either unlock your Default collection for poetry (unjustifiedly), or your GUI session gets effectively locked up, because any time you hit Cancel on the keyring unlock dialog, poetry immediately pops up another, and this dialog grabs the keyboard -- you cannot even switch to a character VT, for killing poetry; you have to log in via ssh for that. This issue is not visible to users who don't use "keyring" (GNOME or otherwise). For those who do, work around the problem by selecting the "null" keyring back-end, in the environment of every poetry invocation. Note: I have not regression-tested the workaround in a desktop environment where "keyring" is unavailable to begin with. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-10-15 12:07:00 +02:00
Botond Dénes	f93abebbb9	Merge 'Sanitize compaction manager API endpoints' from Pavel Emelyanov Endpoints are registered next to the service they use, and the unregistration deferred action is created right after it. When registered, the service in question is passed as argument and then captured by enpoints lambdas. This makes sure that service is not used by endpoints after being stopped. That's not quite the case for compaction manager. Its endpoints can be registered in several places, and compaction_manager "function" is not unregistered on stop. This patch fixes some of this misbehavior, in particular: - adds unregistration of compaction_manager API function - uses sharded<compaction_manager>& argument in endpoints instead of ctx.db.local().get_compaction_manager() chain - moves some endpoints from storage_service.cc to compaction_manager.cc Closes scylladb/scylladb#20962 * github.com:scylladb/scylladb: api: Use captured compaction_manager in get_cm_stats() helper api: Use captured compaction_manager in endpoints api: Add sharded<compaction_manager> argument to compaction_manager API reg/unreg api: Move some endpoints from storage_service.cc to compaction_manager.cc api: Unset compaction_manager endpoints api: Use shorter registration method for compaction_manager function	2024-10-15 10:20:58 +03:00
Daniel Reis	28a265ccd8	docs: fix redirect from cert-based auth to security/enable-auth page Closes scylladb/scylladb#19943	2024-10-15 09:29:05 +03:00
Kefu Chai	a82706eb8f	Update seastar submodule * seastar 3c9c2696...abd20efd (44): > Revert "build: enable Seastar to build shared and static libs in a single build" > dns: Support c-ares before 1.22 > build: improve c-ares version extraction method > Minor typos fix in doc: reference_wrapper.hh > build: enable Seastar to build shared and static libs in a single build > build: include -fno-semantic-interposition in CXXFLAGS > loop: add Sentinel iterator support to parallel_for_each() > dns: use ARES_LIB_INIT_NONE instead of a magic number > dns: use struct typedef for `_channel` > doc/testing.md: explain seastar + boost test colocation > build: extract c-ares version from header file > dns: replace deprecated ares_process() with ares_process_fd() > build: do not support c-ares >= 1.33 > http: fix indentation > http: Add non-owning `make_request` to http client > treewide: replace boost::irange with std::views::iota where possible > Added unit test for "http_content_length_data_sink_impl" > sharded.hh: migrate to concepts > file, scheduling: remove non-unified I/O and CPU scheduling > http: Add more HTTP response codes > http: add constness to `response_line` > http: refactor response_line to use `seastar::format` > httpd/file_handler: Always close stream > build: add compiler and C++ standard compatibility checks > rpc: rpc_types: replace boost::any with std::any > tls: drop dependency on boost::any > rpc: drop unnecessaty includes to boost libraries > rpc: compressor factory: deinline some boost-using functions > sharded: replace boost ranges with <ranges> > scheduling_specific: drop dependency on boost range adaptors > prefetch: drop dependency on boost::mpl > resource: drop unused dependency on boost::any > smp: drop dependency on boost ranges > reactor: remove unnecessary boost includes > execution_stage: remove unnecessary boost includes > sharded.hh: add invoke_on variant for a shard range > shared_ptr: remove deprecated lw_shared_ptr assignment operator > seastar-addr2line: add --debug arg > addr2line: add type checking > warnings: fix unused result warnings > thread_pool: fix includes > signal: remove trailing spaces > tests/unit: chmod -x signal_test.cc > iostream/http: Fix output_stream::write(temporary_buffer) overload Closes scylladb/scylladb#21109	2024-10-15 09:09:29 +03:00
Tomasz Grabiec	3e438d23e1	Merge 'Check system.tablets update before putting it into the table' from Pavel Emelyanov Having tablet metadata with more than 1 pending replica will prevent this metadata from being (re)loaded due to sanity check on load. This patch fails the operation which tries to save the wrong metadata with a similar sanity check. For that, changes submitted to raft are validated, and if it's topology_change that affects system.tablets, the new "replicas" and "new_replicas" values are checked similarly to how they will be on (re)load. fixes #20043 Closes scylladb/scylladb#21020 * github.com:scylladb/scylladb: tablets: Validate system.tablets update group0_client: Introduce change validation group0_client: Add shared_token_metadata dependency	2024-10-15 00:38:59 +02:00
Piotr Smaron	3969ffb39f	test: fix flaky `test_multidc_alter_tablets_rf` The testcase is flaky due to a known python driver issue: https://github.com/scylladb/python-driver/issues/317. This issue causes the `CREATE KEYSPACE` statement to be sometimes executed twice in a row, and the 2nd CREATE statement causes the test to fail. In order to work around it, it's enough to add `if not exists` when creating a ks. Fixes: scylladb/scylladb#21034 Needs to be backported to all 6.x branches, as the PR introducing this flakiness is backported to every 6.x branch. Closes scylladb/scylladb#21056	2024-10-14 16:18:44 +02:00
Avi Kivity	c286ddab38	test: lib: rest_client: use 'http' scheme even when connecting via a unix socket aiohttp 3.10.5 complains when 'unix+http' is used for a unix-domain socket. USe 'http', which work with 3.10.5 and the toolchain's 3.9.5. Closes scylladb/scylladb#21080	2024-10-14 15:32:56 +02:00
Piotr Dulikowski	48d75818fd	SCYLLA-VERSION-GEN: correct the logic for skipping SCYLLA--FILE The SCYLLA-VERSION-GEN file skips updating the SCYLLA--FILE files if the commit hash from SCYLLA-RELEASE-FILE is the same. The original reason for this was to prevent the date in the version string from changing if multiple modes are built across midnight (scylladb/scylla-pkg#826). However - intentionally or not - it serves another purpose: it prevents an infinite loop in the build process. If the build.ninja file needs to be rebuilt, the configure.py script unconditionally calls ./SCYLLA-VERSION-GEN. On the other hand, if one of the SCYLLA-*-FILE files is updated then this triggers rebuild of build.ninja. Apparently, this is sufficient for ninja to enter an infinite loop. However, the check assumes that the RELEASE is in the format <build identifier>.<date>.<commit hash> and assumes that none of the components have a dot inside - otherwise it breaks and just works incorrectly. Specifically, when building a private version, it is recommended to set the build identifier to `count.yourname`. Previously, before `85219e9`, this problem wasn't noticed most likely because reconfigure process was broken and stopped overwriting the build.ninja file after the first iteration. Fix the problem by fixing the logic that extracts the commit hash - instead of looking at the third dot-separated field counting from the left side, look at the last field. Fixes: scylladb/scylladb#21027 Closes scylladb/scylladb#21049	2024-10-14 13:49:15 +03:00
Calle Wilund	8eaf00ff11	test::topology: Add test for TLS upgrade and downgrade of internode encryption Test a rolling upgrade of cluster while active. Note: This is a unit test version of dtest test. Has the big drawback of not being able to use cassandra-stress to work and verify the cluster and results Test moves from none to all to none encryption while writing and then checking written data.	2024-10-13 23:54:06 +00:00
Calle Wilund	a557f699a2	docs: Add internode_encryption=transitional documentation Describing upgrading cluster(s) without downtime.	2024-10-13 23:54:06 +00:00
Calle Wilund	390b9759b6	messaging_service: Add "transitional" internode encryptipn mode Fixes #18903 Adds a "transitional" internode encryption mode, under which all _outgoing_ RPC connections will use TLS, but we will still accept any incoming non-tls connection. This allows an operator to perform a move to TLS RPC without cluster downtime: 1. For each server, add certificate etc options to server_encryption_options + internode_encryption=none + set ssl_storage_port + restart (rolling) 2. For each server, set internode_encryption=transitional + RR 3. For each server, set internode_encryption=all + RR	2024-10-13 23:54:06 +00:00
Calle Wilund	503a71f9b8	messaging_service: Create TLS connector even if internode_enc=none when certs set Refs #18903 If ssl_storage_port is non-zero _and_ we have specified actual certificates are set/exists, create TLS connector for RPC regardless of whether internode encryption is enables. I.e. potentially unused. For transitioning cluster to TLS.	2024-10-13 23:54:05 +00:00
Avi Kivity	db14a01901	Merge 'Use table id as system.sstables partition key' from Pavel Emelyanov The system.sstables (a.k.a. sstables registry) primary key is "string location" as partition key and "uuid generation" as clustering one. The "location" part was taken from table.config.datadir value which, in turn, a string containing path to on-disk files if the table was located locally, e.g. /var/lib/scylla/data/ks/cf-abc123 one. Recently [1] the datadir was moved from table config onto storage options, but this string is still used as registry key. Other than being owned by a table with ID, sstables are accessed by restore-from-object-storage code [2]. To make it work, both storage driver and sstable_directory helper class maintain two formats of object prefixes for sstables components. For S3-backed sstables having a record in registry, the path used is s3://bucket/generation/component. For restore code there are user-provided prefixes that do not match the aforementioned pattern. The selection between those two is now made by checking sstable state, which is not obvious and may cause troubles for tiered storage driver. This patch changes the registry schema so that partition key becomes "uuid owner" and is set to be table.id() value. This is to stop using the local path by S3 backed sstables. Also this change makes it possible for storage driver and sstable directory to rely on the storage options only to tell different bucket prefixes formats from each other. As a side effect, the make_s3_object_name() helper, that generates the proper object name, becomes explicit for restore-from-S3 usage. Now it relies on the sstable::filename() calling this->prefix() behind the scenes and the latter to return the user-provided prefix, which is pretty fragile construction. No need to backport (and it's not going to be easy to do it), storage options feature is still experimental Refs #20675 [1] Refs #20305 [2] Closes scylladb/scylladb#20998 * github.com:scylladb/scylladb: sstables: Flatten S3 object name making sstable_directory: Flatten directory lister creation treewide: Rename sstable registry location field to be owner system_keyspace: Change sstables registry partition key type sstables: Keep location variant on s3 backend too storage_options: Use variant on S3 options sstables: Split sstable::filename() helper sstables: Add s3_storage::owner() helper	2024-10-13 20:08:43 +03:00
Kefu Chai	7d2d44883b	install.sh: install seastar/scripts/addr2line.py as well seastar extracted `addr2line` python module out back in e078d7877273e4a6698071dc10902945f175e8bc. but `install.sh` was not updated accordingly. it still installs `seastar-addr2line` without installing its new dependency. this leaves us with a broken `seastar-addr2line` in the relocatable tarball. ```console $ /opt/scylladb/scripts/seastar-addr2line Traceback (most recent call last): File "/opt/scylladb/scripts/libexec/seastar-addr2line", line 26, in <module> from addr2line import BacktraceResolver ModuleNotFoundError: No module named 'addr2line' ``` in this change, we redistribute `addr2line.py` as well. this should address the issue above. Fixes scylladb/scylladb#21077 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21078	2024-10-13 19:35:14 +03:00
Kefu Chai	519b4a2934	utils/s3: include used header when building the tree with clang-19 and libstdc++ shipped along with GCC 14.2.1, we have ``` clang++ -MD -MT build/release/utils/s3/aws_error.o -MF build/release/utils/s3/aws_error.o.d -std=c++23 -I/home/kefu/dev/scylladb/master/seastar/include -I/home/kefu/dev/scylladb/master/build/release/seastar/gen/include -Werror=unused-result -DSEASTAR_API_LEVEL=7 -DSEASTAR_SSTRING -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_LOGGER_TYPE_STDOUT -DFMT_SHARED -I/usr/include/p11-kit-1 -DWITH_GZFILEOP -ffile-prefix-map=/home/kefu/dev/scylladb/master=. -march=westmere -ffunction-sections -fdata-sections -O3 -mllvm -inline-threshold=2500 -fno-slp-vectorize -DSCYLLA_BUILD_MODE=release -g -gz -Xclang -fexperimental-assignment-tracking=disabled -iquote. -iquote build/release/gen -std=gnu++23 -ffile-prefix-map=/home/kefu/dev/scylladb/master=. -march=westmere -DBOOST_ALL_DYN_LINK -fvisibility=hidden -isystem abseil -Wall -Werror -Wextra -Wimplicit-fallthrough -Wno-mismatched-tags -Wno-c++11-narrowing -Wno-overloaded-virtual -Wno-unused-parameter -Wno-unsupported-friend -Wno-missing-field-initializers -Wno-deprecated-copy -Wno-psabi -Wno-error=deprecated-declarations -DXXH_PRIVATE_API -DSEASTAR_TESTING_MAIN -c -o build/release/utils/s3/aws_error.o utils/s3/aws_error.cc utils/s3/aws_error.cc:33:21: error: no member named 'make_unique' in namespace 'std' 33 \| auto doc = std::make_unique<rapidxml::xml_document<>>(); \| ~~~~~^ utils/s3/aws_error.cc:33:57: error: expected '(' for function-style cast or type construction 33 \| auto doc = std::make_unique<rapidxml::xml_document<>>(); \| ~~~~~~~~~~~~~~~~~~~~~~~~^ utils/s3/aws_error.cc:33:59: error: expected expression 33 \| auto doc = std::make_unique<rapidxml::xml_document<>>(); \| ^ 3 errors generated. ninja: build stopped: subcommand failed. ``` in order to address the build failure, let's include the used header. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21064	2024-10-13 18:32:34 +03:00
Patryk Jędrzejczak	18d3a6480d	test: test_read_required_hosts: run with the raft-based topology When we made the raft-based topology mandatory, all boost test tests started using it. Then, `test_read_required_hosts` started failing. We left investigating it for later and started running it with `force-gossip-topology-changes` to make it pass. Currently, the test doesn't fail with the raft-based topology anymore. Hence, we remove the FIXME and run the test with a normal config. We don't know when and why the test stopped failing. Investigating it wouldn't be easy, since we don't even know why it failed in the first place. We suspect that there was some bug that is now fixed. This patch only fixes a test, there is no need to backport it. Fixes scylladb/scylladb#18463 Closes scylladb/scylladb#20960	2024-10-11 17:01:20 +02:00
Kamil Braun	96070bb5b3	Merge 'storage_proxy: Add conditions checking to avoid UB in speculating read executors.' from Sergey Zolotukhin During the investigation of scylladb/scylladb#20282, it was discovered that implementations of speculating read executors have undefined behavior when called with an incorrect number of read replicas. This PR introduces two levels of condition checking: - Condition checking in speculating read executors for the number of replicas. - Checking the consistency of the Effective Replication Map in filter_for_query(): the map is considered incorrect if the list of replicas contains a node from a data center whose replication factor is 0. Please note: This PR does not fix the issue found in scylladb/scylladb#20282; it only adds condition checks to prevent undefined behavior in cases of inconsistent inputs. Refs scylladb/scylladb#20625 As this issue applies to the releases versions and can affect clients, we need backports to 6.0, 6.1, 6.2. Closes scylladb/scylladb#20851 * github.com:scylladb/scylladb: Add conditions checking for get_read_executor Avoid an extra call to block_for in db::filter_for_query. Improve code readability in consistency_level.cc and storage_proxy.cc tools: Add build_info header with functions providing build type information tests: Add tests for alter table with RF=1 to RF=0	2024-10-11 15:02:02 +02:00
Paweł Zakrzewski	900a6706b8	test: test_restart_cluster: create the test The purpose of this test that the cluster is able to boot up again after a full cluster shutdown, thus exhibiting no issues when connecting to raft group 0 that is larger than one.	2024-10-11 13:25:07 +02:00
Paweł Zakrzewski	7008b71acc	auth: standard_role_manager allows awaiting superuser creation This change implements the ability to await superuser creation in the function ensure_superuser_is_created(). This means that Scylla will not be serving CQL connections until the superuser is created. Fixes #10481	2024-10-11 13:25:07 +02:00
Paweł Zakrzewski	04fc82620b	auth: coroutinize the standard_role_manager start() function This change is a preparation for the next change. Moving to coroutines makes the code more readable and easier to process.	2024-10-11 13:25:07 +02:00
Paweł Zakrzewski	f525d4b0c1	auth: don't start server until the superuser is created This change reorganizes the way standard_role_manager startup is handled: now the future returned by its start() function can be used to determine when startup has finished. We use this future to ensure the startup is finished prior to starting the CQL server. Some clusters are created without auth, and auth is added later. The first node to recognize that auth is needed must create the superuser. Currently this is always on restart, but if we were to ever make it LiveUpdate then it would not be on restart. This suggests that we don't really need to wait during restart. This is a preparatory commit, laying ground for implementation of a start() function that waits for the superuser to be created. The default implementation returns a ready future, which makes no change in the code behavior.	2024-10-11 13:25:07 +02:00
Pavel Emelyanov	a7042d66e3	sstables: Flatten S3 object name making The s3_storage backend driver has a method that generates object path within the bucket. Depending on options alternative it picks one of two formats: - for string prefix, it uses it implicitly via sstable::filename() call that calls storage->prefix() which, in turn, returns prefix value - for registry-backed sstables, the /bucket/generation/component path is generated This patch bruses this place up. Similarly to previous patch, this change also makes the selection based on the location alternative, not on the sstable state. As well it's idempotent change, as S3 sstables with 'upload' state only appear when restoring from object store, and in this case the string location is in use. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 14:11:28 +03:00
Pavel Emelyanov	8d5537a439	sstable_directory: Flatten directory lister creation After previous patchin, the way components lister is created for S3 storage options became quite hairy. This patch brushes things up to be easier to read. The only "functional" change here, is that selection between registry lister and S3 lister is made based on options' location held alternative, not on the sstable state value. That's in fact idempotent change, the only caller that provides string location on options is the "restore from object store" code that also sets state to be 'upload'. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 14:11:28 +03:00
Pavel Emelyanov	031893259a	treewide: Rename sstable registry location field to be owner This is sort of continuation of the previous patch. The partition key in the registry is now table_id, not string, and is better called "owner", not "location". This patch is s/location/owner/ over specific places that include field name in the schema, argument names in registry maintenance classes and tests accessing the selected row fields by name. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 14:11:28 +03:00
Pavel Emelyanov	3315e3a2a9	system_keyspace: Change sstables registry partition key type Today, the system.sstables schema uses string as partition key. Callers, in turn, use table's datadir value to reference entries in it. That's wrong, S3-backed sstables don't have any local paths to work with. The table's ID is better in this role. This patch only changes the field type to be table_id and fixes the callers to provide one. In particular, see init_table_storage() change -- instead of generating a datadir string, it sets table.id() as the options' location. Other fixed places are tests. Internally, this id value is propagated via s3_storage::owner() method, that's fixed as well. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 13:48:09 +03:00
pehala	a2f9136e36	test.py: Use "python -m pytest" for pytest invocation for PythonTest Enables debugging inside pytest subprocesses as well. It seems that pydev automatically attaches itself also to all python subprocesses. Since we used to call "pytest" wrapper it was deemed a different program, and we could not debug individual tests. Closes scylladb/scylladb#21050	2024-10-11 13:38:47 +03:00
Pavel Emelyanov	bb13b7bf72	sstables: Keep location variant on s3 backend too Previous patch put variant<string, table_id> as location of S3 options. This patch makes the S3 sstables backend driver keep variant as sstable location. As with the previous patch, driver only keeps variant, but continues using its string alternative internally. This will be changed later on. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 13:09:47 +03:00
Pavel Emelyanov	1181b6b082	storage_options: Use variant on S3 options Describing S3 storage for an sstables nowadays has two options -- via sstables registry entry and by using the direct prefix string. The former is used when putting a keyspace on S3. In this case each sstable has the corresponding entry in the system.sstables table. The latter is used by "restore from object storage" code. In that case, sstables don't have entries in the registry, but are accessed by a specific S3 object path. This patch reflects this difference by making s3_options::location be variant of string prefix and table_id owner. The owner needs more explanation, here it is. Today, the system.sstables schema defines partition key to be "string location" and clustering key to be "UUID generation". The partition key is table's datadir string, but it's wrong to use it this way. Next patches will change the partition key to be table's ID (there's table_id type for it), and before doing it storage options must be prepared to carry it onboard. This patch does it, but the table_id alternative of the location is still unused, the rest of the code keeps using the string location to reference a row in the registry table. Next patches will eventually make use of the table_id value. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 13:04:52 +03:00
Kamil Braun	4d99cd2055	Merge 'raft: fast tombstone GC for group0-managed tables' from Emil Maskovsky Add the gossip state for broadcasting the nodes state_id. Implemented the Group0 state broadcaster (based on the gossip) that will broadcast the state id of each node and check the minimal state id for the tombstone GC. When there is a change in the tombstone GC minimal state id, the state broadcaster will update the tombstone GC time for the group0-managed tables. The main component of the change is the newly added `group0_state_id_handler` that keeps track, broadcasts and receives the last group0 state_ids across all nodes and sets the tombstone GC deletion time accordingly: * on each group0 change applied, the state_id handler broadcasts the state_id as a gossip state (only if the value has changed) * the handler checks for the node state ids every refresh period (configurable, 1h by default) * on every check, the handler figures out the lowest state_id (timeuuid), which is state_id that all of the nodes already have * the timestamp of this minimum state_id is then used to set the tombstone GC deletion time * the tombstone GC calculation then uses that deletion time to provide the GC time back to the callers, e.g. when doing the compaction * (as the time for tombstone GC calculation has the 1s granularity we actually deduce 1s from the determined timestamp, because it can happen that there were some newer mutations received in the same second that were not distributed across the nodes yet) This change introduces a new flag to the static schema descriptor (`is_group0_table`) that is being checked for this newly added mode in the tombstone GC. We also add a check (in non-release builds only) on every group0 modification that the table has this flag set. The group0 tombstone GC handling is similar to the "repair" tombstone GC mode in a sense (that the tombstone GC time is determined according to a reconciliation action), however it is not explicitly visible to (nor editable by) the user. And also the tombstone GC calculation is much simpler than the "repair" mode calculation - for example, we always use the whole range (as opposed to the "repair" mode that can have specific repair times set for specific ranges). We use the group0 configuration to determine the set of nodes (both current and previous in case of joint configuration) - we need to make sure that we account for all the group0 nodes (if any node didn't provide the state_id yet, the current check round will be skipped, i.e. no GC will be done until all known nodes provide their state_id timestamp value). Also note that the group0 state_id handling works on all nodes independently, i.e. each node might have its own (possibly different) state depending on the gossip application state propagation. This is however not a problem, as some nodes might be behind, but they will catch up eventually, and this solution has the benefit of being distributed (as opposed to having a central point to handle the state, like for example the topology coordinator that has been considered in the early stages of the design). Fixes: scylladb/scylla#15607 New feature, should not be backported. Closes scylladb/scylladb#20394 * github.com:scylladb/scylladb: raft: add the check for the group0 tables raft: fast tombstone GC for group0-managed tables tombstone_gc: refactor the repair map raft: flag the group0-managed tables gossip: broadcast the group0 state id raft/test: add test for the group0 tombstone GC treewide: code cleanup and refactoring	2024-10-11 11:52:27 +02:00
Pavel Emelyanov	ba97072709	sstables: Split sstable::filename() helper To have the filename(type, prefix) one, next patches will provide prefix on their own, to avoid storage->prefix() call. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 12:47:13 +03:00
Pavel Emelyanov	6f9cb51259	sstables: Add s3_storage::owner() helper This driver uses sstring _location as part of the lookup key in the sstables registry. Next patches will need to change that and put more checks on the registry access, so introduce a helper method beforehand. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-11 12:47:12 +03:00
Sergey Zolotukhin	c373edab2d	Add conditions checking for get_read_executor During the investigation of scylladb/scylladb#20282, it was discovered that implementations of speculating read executors have undefined behavior when called with an incorrect number of read replicas. This PR introduces two levels of condition checking: - Condition checking in speculating read executors for the number of replicas. - Checking the consistency of the Effective Replication Map in get_endpoints_for_reading(): the map is considered incorrect the number of read replica nodes is higher than replication factor. The check is applied only when built in non release mode. Please note: This PR does not fix the issue found in scylladb/scylladb#20282; it only adds condition checks to prevent undefined behavior in cases of inconsistent inputs. Refs scylladb/scylladb#20625	2024-10-11 09:38:25 +02:00
Sergey Zolotukhin	8db6d6bd57	Avoid an extra call to block_for in db::filter_for_query.	2024-10-11 09:38:25 +02:00
Sergey Zolotukhin	ad93cf5753	Improve code readability in consistency_level.cc and storage_proxy.cc Add const correctness and rename some variables to improve code readability.	2024-10-11 09:38:25 +02:00
Sergey Zolotukhin	ae23d42889	tools: Add build_info header with functions providing build type information A new header provides `constexpr` functions to retrieve build type information: `get_build_type()`, `is_release_build()`, and `is_debug_build()`. These functions are useful when adding changes that should be enabled at compile time only for specific build types.	2024-10-11 09:38:24 +02:00
Sergey Zolotukhin	132358dc92	tests: Add tests for alter table with RF=1 to RF=0 Adding Vnodes and Tablets tests for alter keyspace operation that decreases replication factor from 1 to 0 for one of two data centers. Tablet version fails due to issue described in scylladb/scylladb#20625. Test for scylladb/scylladb#20625	2024-10-11 09:38:24 +02:00
Pavel Emelyanov	77eb9ddb0f	sstable_set: Reserve vector of readers When generating readers for the set of sstables, the end size of this vector is known in advance and its storage can be reserved. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21055	2024-10-11 09:56:17 +03:00
Pavel Emelyanov	551da72492	api: Use captured database, not the one from ctx Continuation of the previous patch -- not commitlog-related endpoints can use provided database reference, that was captured from main. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 17:53:30 +03:00
Pavel Emelyanov	74f7071db8	api: Pass sharded<database> to commitlog endpoints registration This is to make registered enpoints with with the database without grabbing one from ctx. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 17:52:56 +03:00
Pavel Emelyanov	14ab6d2615	api: Move commitlog-related from storage_service.cc It registers itself in /storage_service function, but works with commitlog, so should be located next to commitlog endpoints. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 17:52:25 +03:00
Pavel Emelyanov	ba73704774	api: Unset commitlog API endpoints Most of other set_...()-s has the unset_...() scheduled right afterwards, so here's one for set_server_commitlog(). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 17:50:37 +03:00
Pavel Emelyanov	44ec6d36f3	api: Extract set_server_commitlog() from set_server_done() The latter collects a bunch of endpoints including commitlog ones. Extract it as snandalone call in main. It's currently not located next to "commitlog server" as it should, because there's no standalone commitlog service in main. It will be addressed as a followup together with other endpoints that work with sharded<database>. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 17:49:10 +03:00
Pavel Emelyanov	1863ccd900	tablets: Validate system.tablets update Implement change validation for raft topology_change command. For now the only check is that the "pending replicas" contains at most one entry. The check mirrors similar one in `process_one_row` function. If not passed, this prevents system.tablets from being updated with the mutation(s) that will not be loaded later. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 12:39:58 +03:00
Pavel Emelyanov	e5bf376cbc	group0_client: Introduce change validation Add validate_change() methods (well, a template and an overload) that are called by prepare_command() and are supposed to validate the proposed change before it hits persistent storage Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 12:31:52 +03:00
Pavel Emelyanov	f09fe4f351	group0_client: Add shared_token_metadata dependency It will be needed later to get tablet_metadata from. The dependency is "OK", shared_token_metadata is low-level sharded service. Client already references db::system_keyspace, which in turn references replica::database which, finally, references token_metadata Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-10 12:27:46 +03:00
Botond Dénes	86fd9ce8fd	schema/schema: break circular dependency with replica::database The schema module (everything in schema/) is supposed to be towards the leafs in the ScyllaDB inter-module dependency graph. In other words, it should not depend on many other modules. On the other hand, almost the entire codebase depends on the schema module itself. Currently there is a circular dependency between schema and replica::database, as the latter is a required argument for schema::describe(). This is bad, not just because of the dependency mess it introduces, but also because now schema::describe() can only be used by code which has a reference to the database handy. This patch breaks this circular dependency, by introducing the schema_describe_helper interface and providing an implementation for it in database.hh. There is another circular dependency: schema <-> replica::table. This is not addressed by this patch. Closes scylladb/scylladb#20893	2024-10-10 10:07:26 +03:00
Botond Dénes	81423e8e76	Merge 'repair: Fix stall in repair_get_row_diff_with_rpc_stream_process_op_slow_path' from Asias He Use clear_gently to avoid the following stalls. ``` ~frozen_mutation_fragment at ././frozen_mutation.hh:268 std::destroy_at<frozen_mutation_fragment> at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/stl_construct.h:88 std::allocator_traits<std::allocator<std::_List_node<frozen_mutation_fragment> > >::destroy<frozen_mutation_fragment> at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/alloc_traits.h:537 std::__cxx11::_List_base<frozen_mutation_fragment, std::allocator<frozen_mutation_fragment> >::_M_clear at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/list.tcc:77 ~_List_base at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/stl_list.h:499 ~partition_key_and_mutation_fragments at ././repair/repair.hh:298 ~repair_row_on_wire_with_cmd at ././repair/repair.hh:335 operator() at ./repair/row_level.cc:1881 ``` Fixes #21016 Performance improvement only. No backport. Closes scylladb/scylladb#21017 * github.com:scylladb/scylladb: repair: Fix stall in repair_get_row_diff_with_rpc_stream_process_op_slow_path repair: Add clear_gently for partition_key_and_mutation_fragments	2024-10-10 09:27:27 +03:00
Benny Halevy	3a12ad96c7	sstables: scylla_metadata: add sstable identifier Keep a copy of the sstable uuid generation in a new scylla_metadata sstable_identifier attribute. If the SSTable happens to have a numerical generation just create a new time-uuid and log a message about that. Dump this new attribute in scylla sstable dump tool. And add a unit test to verify that the written (and then loaded) sstable identifier matches the sstable's generation. The motivatrion for this change stems from backup deduplication. In essence, an sstable may already have been backed up in a previous snapshot, and we don't want to abck it up again if it's already present on external storage. Today this is based on rclone that compares files checksums, but once scylla will backup the sstables using the native object-storage stack (#19890), we would like to use the sstable globally-unique identifier for deduplication. Although the uuid-generation is encoded in the sstable path, the latter may change, e.g. due to intra-node migration, so keep a copy of the original unique identifier in scylla-metadata, and that attribute would survive file-based or intra-node migrations. Fixes scylladb/scylladb#20459 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#21002	2024-10-10 08:52:46 +03:00
Avi Kivity	b66479ea98	Merge 'compaction: fix potential data resurrection with file-based migration' from Ferenc Szili When tablets are migrated with file-based streaming, we can have a situation where a tombstone is garbage collected before the data it shadows lands. For instance, if we have a tablet replica with 3 sstables: 1. sstable containing an expired tombstone 2. sstable with additional data 3. sstable containing data which is shadowed by the expired tombstone in sstable 1 If this tablet is migrated, and the sstables are streamed in the order listed above, the first two sstables can be compacted before the third sstable arrives. In that case, the expired tombstone will be garbage collected, and data in the third sstable will be resurrected after it arrives to the pending replica. This change fixes this problem by disabling tombstone garbage collection for pending replicas. This fixes a problem in Enterprise, but the change is in OSS in order to have as few differences between OSS and Enterprise and to have a common infrastructure for disabling tombstone GC on pending replicas. This change has to be backported to all active versions: 6.0, 6.1 and 6.2, as well as Enterprise 2024.2 Closes scylladb/scylladb#20788 * github.com:scylladb/scylladb: test: test tombstone GC disabled on pending replica tablet_storage_group_manager: update tombstone_gc_enabled in compaction group database::table: add tombstone_gc_enabled(locator::tablet_id)	2024-10-09 21:49:49 +03:00
Avi Kivity	bb1867c7c7	Merge 'sstables: Add digest checking in the validation path of the sstable layer' from Nikos Dragazis This PR builds upon the PR for checksum validation (#20207) to further enhance scrub's corruption detection capabilities by validating digests as well. The digest (full checksum) is the checksum over the entire data, as opposed to per-chunk checksums which apply to individual chunks. Until now, digests were not examined on any code paths. This PR integrates digest checking into the compressed/checksummed data sources as an optional feature and enables it only through the validation path of the sstable layer (`sstable::validate()`). The validation path is used by the following tools: * scrub in validate mode * `sstable validate` All other reads, including normal user reads, are unaffected by this change. The PR consists of: * Extensions to the compressed and checksummed data sources to support digest checking. The data sources receive the expected digest as a parameter and calculate the actual digest incrementally across multiple get() calls. The check happens on the get() call that reaches EOF and results to an exception if the digest is invalid. A digest check requires reading the whole file range. Therefore, a partial read or skip() is treated as an internal error. * A new shareable digest component loaded on demand by the validation code. No lifecycle management. * Grouping of old scrub/validate tests for compressed and uncompressed SSTables to reduce code duplication. * scrub/validate tests for SSTables with valid checksums but invalid digests, and SSTables with no digests at all. * scrub/validate tests with 3.x Cassandra SSTables to ensure compatibility. Refs #19058. New feature, no backport is needed. Closes scylladb/scylladb#20720 * github.com:scylladb/scylladb: test: Test scrub/validate with SSTables from Cassandra compaction: Make quarantine optional for perform_sstable_scrub() test: Make random schema optional in scrub_test_framework test: Add tests for invalid digests test: Merge scrub/validate tests for compressed and uncompressed cases sstables: Verify digests on validation path sstables: Check if digest component exists sstables: Add digest in the SSTable components sstables: Add digest check in compressed data source sstables: Add digest check in checksummed data source	2024-10-09 21:33:08 +03:00
Benny Halevy	d34878e96c	view: check_needs_view_update_path: get token_metadata_ptr check_needs_view_update_path is async and might yield so the token_metadata reference passed to it must be kept alive throughout the call. Fixes scylladb/scylladb#20979 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#20980	2024-10-09 20:56:21 +03:00
Nadav Har'El	a1999cd5d5	cql-pytest: fix run-cassandra on systems with default Java 8 The test/cql-ptest/run-cassandra prefers to use Java 11 if installed on the system because this is the only version of Java that all modern versions of Cassandra run on (Cassandra 3 and 4 can run on Java 8 and 11, Cassandra 5 can run on Java 11 and 17). However, in our search order we tried the "java" in the user's path first, before trying Java 11. This means that if the user for some reason had the ancient Java 8 (which is now a decade old) as his default "java" got that, instead of Java 11, and couldn't run Cassandra 5. While at it, update the comments to reflect the new reality that Cassandra 5 needs Java 17 or 11 - not 11 or 8 as the older Cassandra. We should eventually change the code logic as well (searching for versions that depend on the Cassandra version - not always Java 8 and 11), but let's do it later. This patch already fixes a real bug for developers that did install Java 11 but their default "java" pointed to Java 8. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#21001	2024-10-09 20:51:56 +03:00
David Garcia	2247bdbc8c	docs: Fix confgroup links It was not possible to link to configuration parameters groups in docs/reference/configuration-parameters.rst if they contained a space. Closes scylladb/scylladb#21018	2024-10-09 20:16:15 +03:00
Pavel Emelyanov	3dcf3d65d7	replica: Use substract_sets() helper The process_one_row() evaluates pending_replica by subtracting replicas from new_replicas. There's a convenience helper for that. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#21019	2024-10-09 20:02:16 +03:00
Gleb Natapov	f7e7e61fa7	raft: add more information to start_read_barrier error Add requester into to the error about requester being out of config. Also fix a typo while we are at it. Message-ID: <ZwVDtOty2cWy3vqD@scylladb.com>	2024-10-09 16:24:34 +02:00
Pavel Emelyanov	7163fbcef5	Merge 'utils: replace dependency on boost ranges with <ranges>' from Avi Kivity To avoid depending on two similar libraries (boost ranges and std \<ranges), replace uses of the former with the latter. This series tackles the utils/ directory. Code cleanup, no backport. Closes scylladb/scylladb#20997 * github.com:scylladb/scylladb: utils: logalloc: replace boost with std utils: lsa: chunked_managed_vector: replace boost with std utils: config_file: replace boost with std utils: loading_cache: replace boost with std utils: fragment_range: replace boost with std utils: error_injector: replace boost with std utils: crc: replace boost for_each with built-in range for utils: class_registrator: replace boost with std utils: chunked_vector: replace boost with std utils: observable: replace boost with std	2024-10-09 16:04:48 +03:00
Botond Dénes	3e468608e7	Merge 'Collect sstables on boot from all datadirs (and don't collect from S3 twice)' from Pavel Emelyanov There's a long-pending issue in distributed loader. When it populates sstables on boot it loops over table.config.all_datadirs, but ignores the loop cursor (the datadir itslef), instead loading sstables from table.config.dir, which is 0th element of all_datadirs. There's a test for that, but it's also broken. Effectively collection happens from table.config.dir several times. For local sstables that's just wasted work and potentially lost sstables (but nobody seems to configure more than 1 datadir anyway). For S3 sstables it's also wasted work and incorrectness. The fix is for both -- populator and test. The former is to use all_datadirs to construct sstable_directory. To make it happen, creation of sstable_directory now depends on the storage options, the loop is moved into the branch that creates sstable_directory for local storage type. The test fix is to make sure that some sstables in non-default datadir before running population code. Closes scylladb/scylladb#20819 * github.com:scylladb/scylladb: test: Fix test_multiple_data_dirs distributed_loader: Indentation fix after previous patch distributed_loader: Use correct datadir to collect local sstable distributed_loader: Move all-datadirs loop to local storage collecting distributed_loader: Collect table subdirs based on its storage options distributed_loader: Indentation fix after previous patch distributed_loader: Squash loop of collect_subdir into one method distributed_loader: Convert map of directories into a vector distributed_loader: Make start_subdir() method work with directory distributed_loader: Drop local reference variable distributed_loader: Split start_subdir() distributed_loader: Remove allow-offstrategy argument distributed_loader: Make populate() method work with directory distributed_loader: Remove check for sstable_directory presense distributed_loader: Out-line table_populator() methods distributed_loader: Print storage options, not datadir distributed_loader: Print prepared message sstable_directory: Add sstable_state argument ot one of constructors sstable_directory: Add state() method	2024-10-09 14:43:34 +03:00
Michał Chojnowski	c2ba300f1c	reader_concurrency_semaphore: in stats, fix swapped count_resources and memory_resources can_admit_read() returns reason::memory_resources when the permit is queued due to lack of count resources, and it returns reason::count_resources when the permit is queued due to lack of memory resources. It's supposed to be the other way around. This bug is causing the two counts to be swapped in the stat dumps printed to the logs when semaphores time out. Closes scylladb/scylladb#20714	2024-10-09 14:12:01 +03:00
Lakshmi Narayanan Sreethar	69c385f540	compaction: make drain wait for compactions to stop during shutdown During shutdown, the compaction_manager starts stopping ongoing compaction tasks through `really_do_stop()` method as soon as it receives a signal from the abort source. Later, when the database object shuts down, it calls `compaction_manager::drain` to ensure that all compaction tasks have stopped. However, `compaction_manager::drain` is currently implemented in such a way that, during shutdown, it effectively becomes a no-op because the compaction_manager has already initiated the stopping of tasks. As a result the caller assumes that all the compaction tasks have stopped and proceeds to close all the tables. This can lead to race conditions where table closures overlap with compaction tasks that are still running, resulting in exceptions like : ``` exception during mutation write to 127.0.0.1: utils::internal::nested_exception<std::runtime_error> (Could not write mutation system:compaction_history (pk{0010b70d31705e0411efb2edf6467f094c8b}) to commitlog): seastar::gate_closed_exception (gate closed) ``` This commit fixes the issue by updating `compaction_manager::drain` to invoke `stop_ongoing_compactions` even during shutdown to ensure that it waits for the ongoing compaction tasks to complete. The `stop_ongoing_compactions` method will also send a stop request to these tasks before waiting, but the request will be ignored by the tasks as they would have already received one earlier from `really_do_stop()`. Fixes #20197 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com> Closes scylladb/scylladb#20715	2024-10-09 12:08:32 +03:00
Pavel Emelyanov	17ec416178	Merge 'Make sure S3 upload completion parses possible error' from Ernest Zaslavsky fixes #20517 Adds `aws_error` which possibly can contain errors from the S3 response body. Adds to the multipart upload completion a check for possible error and issues a retry if the error is retryable Closes scylladb/scylladb#20518 * github.com:scylladb/scylladb: test: add complete_multipart_upload completion tests code: s3 client error handling code: add response parsing and error handling to the complete_multipart_upload code: Introduce AWS errors parsing	2024-10-09 12:01:27 +03:00
Piotr Smaron	e0c1a51642	cql/tablets: handle MVs in ALTER tablets KEYSPACE ALTERing tablets-enabled KEYSPACES (KS) didn't account for materialized views (MV), and only produced tablets mutations changing tables. With this patch we're producing tablets mutations for both tables and MVs, hence when e.g. we change the replication factor (RF) of a KS, both the tables' RFs and MVs' RFs are updated along with tablets replicas. The `test_tablet_rf_change` testcase has been extended to also verify that MVs' tablets replicas are updated when RF changes. Fixes: #20240 Closes scylladb/scylladb#21007	2024-10-09 10:51:18 +02:00
Pavel Emelyanov	0bc8d0c620	Merge 'utils: unconst: wean away from boost range library' from Avi Kivity As part of the effort to standardize on a single range library, convert the unconst helper and its only user to \<ranges>. The only user, mutation_partitions, happens to use intrusive_btree::iterator as the payload. That iterator wasn't fully conform to iterator requirements, so it's fixed in a preliminary patch. Code cleanup; no backport. Closes scylladb/scylladb#20986 * github.com:scylladb/scylladb: utils/unconst, mutation_partition: switch to ranges utils: intrusive_btree: improve conformity with iterator requirements	2024-10-09 10:06:52 +03:00
Yuao Ma	1cc7821d12	tools: fix typos in the code This patch corrects a minor typo without any functional changes. Signed-off-by: Yuao Ma <c8ef@outlook.com> Closes scylladb/scylladb#20975	2024-10-09 08:18:36 +03:00
Asias He	2d8442f663	repair: Fix stall in repair_get_row_diff_with_rpc_stream_process_op_slow_path Use clear_gently to avoid the following stalls. ``` ~frozen_mutation_fragment at ././frozen_mutation.hh:268 std::destroy_at<frozen_mutation_fragment> at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/stl_construct.h:88 std::allocator_traits<std::allocator<std::_List_node<frozen_mutation_fragment> > >::destroy<frozen_mutation_fragment> at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/alloc_traits.h:537 std::__cxx11::_List_base<frozen_mutation_fragment, std::allocator<frozen_mutation_fragment> >::_M_clear at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/list.tcc:77 ~_List_base at /usr/lib/gcc/x86_64-redhat-linux/11/../../../../include/c++/11/bits/stl_list.h:499 ~partition_key_and_mutation_fragments at ././repair/repair.hh:298 ~repair_row_on_wire_with_cmd at ././repair/repair.hh:335 operator() at ./repair/row_level.cc:1881 ``` Fixes #21016	2024-10-09 09:37:49 +08:00
Asias He	f5f26e1bba	repair: Add clear_gently for partition_key_and_mutation_fragments It is used to clear mutation_fragments to avoid stalls.	2024-10-09 09:31:14 +08:00
Emil Maskovsky	0c9308cf48	raft: add the check for the group0 tables Added the runtime check to ensure that all the tables that are used with the group0 commands are marked as group0 tables.	2024-10-08 21:08:11 +02:00
Emil Maskovsky	a03e98d6e8	raft: fast tombstone GC for group0-managed tables Set the tombstone GC time for group0-managed tables to the minimal state id of the group0 nodes. The check is being done based on a timer, iterating through each node (according to the group0 topology configuration) and taking the minimum across all nodes. This miminum timestamp is then be used to set the tombstone GC time for the tombstone GC of all the group0-managed tables. Fixes: scylladb/scylla#15607	2024-10-08 21:07:30 +02:00
Emil Maskovsky	74bd79bbb3	tombstone_gc: refactor the repair map Move the repair_map definition to the tombstone_gc file where it is mostly being used. Refactor and add the accessors and setters for the group0 tombstone GC time.	2024-10-08 20:53:54 +02:00
Emil Maskovsky	22471410e7	raft: flag the group0-managed tables Add the schema flag to indicate the group0-managed tables. This is to be used to identify and list the group0-managed tables.	2024-10-08 20:53:54 +02:00
Emil Maskovsky	baea9cfa67	gossip: broadcast the group0 state id Implemented the group0 state_id handler (based on the gossip) that will broadcast the group0 state id of each node. This will be used to set the tombstone GC time for the group0 tables.	2024-10-08 20:53:54 +02:00
Emil Maskovsky	fa45fdf5f7	raft/test: add test for the group0 tombstone GC Test that the group0 fast tombstone GC works correctly.	2024-10-08 20:53:54 +02:00
Emil Maskovsky	a840949ea0	treewide: code cleanup and refactoring Fix the clang-tidy warnings, code cleanup and improvements. Applied the clang format to the updated places.	2024-10-08 20:53:54 +02:00
Nadav Har'El	b4df07df71	Merge 'cql3: Print arguments and return type without frozen when describing UDF' from Dawid Mędrek Scylla doesn't allow for the types of arguments or the return type of a UDF to be frozen. As a result, before these changes, create statements produced to restore UDFs as part of `DESCRIBE` statements could not be executed. Fixes scylladb/scylladb#20256 Backport: necessary as the restore process may not work correctly without these changes. The affected versions span from 5.2 to the current master, but we only want to apply the fix to the live versions, so 6.0, 6.1, and 6.2. Closes scylladb/scylladb#20816 * github.com:scylladb/scylladb: cql3/functions/user_function: Print arguments and return type without frozen cql3/functions/user_function: Use fmt to format create statement	2024-10-08 16:05:28 +03:00
Kamil Braun	2d9b8f269f	Merge 'cql: improve validating RF's change in ALTER tablets KS' from Piotr Smaron This patch series fixes a couple of bugs around validating if RF is not changed by too much when performing ALTER tablets KS. RF cannot change by more than 1 in total, because tablets load balancer cannot handle more work at once. Fixes: #20039 Should be backported to 6.0 & 6.1 (wherever tablets feature is present), as this bug may break the cluster. Closes scylladb/scylladb#20208 * github.com:scylladb/scylladb: cql: sum of abs RFs diffs cannot exceed 1 in ALTER tablets KS cql: join new and old KS options in ALTER tablets KS cql: fix validation of ALTERing RFs in tablets KS cql: harden `alter_keyspace_statement.cc::validate_rf_difference` cql: validate RF change for new DCs in ALTER tablets KS cql: extend test_alter_tablet_keyspace_rf cql: refactor test_tablets::test_alter_tablet_keyspace cql: remove unused helper function from test_tablets	2024-10-08 14:33:45 +02:00
Kamil Braun	1b9337bf99	Merge 'Wait for all users of group0 server to complete before destroying it' from Gleb Natapov Group0 server is often used in asynchronous context, but we do not wait for them to complete before destroying the server. We already have shutdown gate for it, so lets use it in those asynch functions. Also make sure to signal group0 abort source if initialization fails. Fixes scylladb/scylladb#20701 Backport to 6.2 since it contains `af83c5e53e` and it made the race easier to hit, so tests became flaky. Closes scylladb/scylladb#20891 * github.com:scylladb/scylladb: group: hold group0 shutdown gate during async operations group0: Stop group0 if node initialization fails	2024-10-08 13:46:54 +02:00
Avi Kivity	48ea51029f	Merge 'time_window_compaction_strategy: estimated_pending_compactions: reestimate compactions rather than using cached value' from Benny Halevy Currently, `estimated_pending_compactions` uses a precalculated value calculated by `update_estimated_compaction_by_tasks`, which, in turn, is called by `get_compaction_candidates`. That means that, if `estimated_pending_compactions` is called, e.g. right after major compaction, it will return an outdated value that was calculated prior to major compaction, and so, it is no longer relevant. Instead, just recalculate the value in `estimated_pending_compactions` and drop `update_estimated_compaction_by_tasks`. * Enhancement, no backport required Closes scylladb/scylladb#20892 * github.com:scylladb/scylladb: test: cql-pytest: test_compaction: add test_compactionstats_after_major_compaction test/cql-pytest: rename test_compaction{_tombstone_gc,} time_window_compaction_strategy: estimated_pending_compactions: reestimate compactions rather than using cached value	2024-10-08 13:29:51 +03:00
Gleb Natapov	d62fbd795b	storage_proxy: make sure there is no end iterator in _live_iterators array storage_proxy::cancellable_write_handlers_list::update_live_iterators assumes that iterators in _live_iterators can be dereferenced, but the code does not make any attempt to make sure this is the case. The iterator can be the end iterator which cannot be dereferenced. The patch makes sure that there is no end iterator in _live_iterators. Fixes scylladb/scylladb#20874 Closes scylladb/scylladb#20977	2024-10-08 13:16:27 +03:00
Avi Kivity	656dc438ab	utils: logalloc: replace boost with std	2024-10-08 12:07:14 +03:00
Avi Kivity	84b25a51f5	utils: lsa: chunked_managed_vector: replace boost with std	2024-10-08 12:03:30 +03:00
Avi Kivity	fa772701be	utils: config_file: replace boost with std	2024-10-08 12:03:15 +03:00
Avi Kivity	b62fadae5f	utils: loading_cache: replace boost with std Unfortunately, the replacement for boost::range::join(), std::views::concat(), is in C++26 (and not implemented in libstdc++ 14). We use array/transform/join to simulate it.	2024-10-08 11:54:34 +03:00
Laszlo Ersek	934b42c6a8	cmake/check_headers: correct typos Commit `efd65aebb2` ("build: cmake: add check-header target", 2023-11-13) introduced three typos: - In "cmake/check_headers.cmake", it checked whether the "parsed_args_GLOB_RECURSE" argument was defined, but then it referenced the same under the wrong name "parsed_args_RECURSIVE". - The above error masked two further typos; namely the duplicate use of "api" and "streaming" each, as targets. With "parsed_args_GLOB_RECURSE" above fixed, CMake now reports these conflicting arguments (target names). They should have been "node_ops" and "sstables", respectively. Correct the typos. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#20992	2024-10-08 09:38:16 +03:00
Dawid Mędrek	8582ed513b	cql3/functions/user_function: Print arguments and return type without frozen Scylla doesn't allow for the types of arguments or the return type to be frozen. As a result, before these changes, create statements produced to restore UDFs as part of `DESCRIBE` statements could not be executed. We fix that and add a reproducer test and another one to verify that the implementation is correct.	2024-10-07 20:53:10 +02:00
Avi Kivity	72a39b84b0	utils: fragment_range: replace boost with std	2024-10-07 21:32:16 +03:00
Avi Kivity	c8a68c4cf7	utils: error_injector: replace boost with std	2024-10-07 21:28:36 +03:00
Avi Kivity	c560686d92	utils: crc: replace boost for_each with built-in range for Simpler.	2024-10-07 21:19:14 +03:00
Avi Kivity	44419fc5ec	utils: class_registrator: replace boost with std	2024-10-07 21:16:03 +03:00
Avi Kivity	adb92a6c16	utils: chunked_vector: replace boost with std	2024-10-07 21:11:23 +03:00
Avi Kivity	b259389a3e	utils: observable: replace boost with std	2024-10-07 21:11:07 +03:00
Nadav Har'El	45ccceb137	alternator: add "dc" and "rack" options to "/localnodes" request Before this patch, the "/localnodes" HTTP request to the Alternator server lists all the live nodes of the current DC. This patch adds two optional parameters to this query: dc: allows to list the live nodes of a specific named DC instead of the current DC of the server. rack: allows to restrict the results to just the nodes belonging to a specific named rack. For both options, if no live node exists in the given dc or rack (in particular, if such a dc or rack doesn't even exist), an empty list is returned - it's not an error. The default, if dc or rack is not specified - remains exactly as it is today - look at the current DC (the one of the node being request), and do not restrict the list to any specific rack. We expect the new options that we added here to be useful for two use cases: 1. A client that knows of some Scylla node (belonging to an unknown DC), but wants to list the nodes in its DC, which it knows by name. 2. A client in a multi-rack DC (e.g., multi-AZ region in AWS) that wants to send requests to nodes in its own rack (which it knows by name), to avoid cross-rack networking costs. Note that in both cases, this requires clients to know the names of DCs and AZs via some out-of-band means. The client can also get a list of DCs and racks using the system.local system table, as the tests included in this patch demonstrate. This patch includes two set of tests for these new options: One in the the single-node test/alternator framework that has a single dc and rack but can still check the case of an unknown dc or rack (in which case an empty list is returned). The second test is in the topology framework, and runs an 8-node cluster with two DCs, two racks, and two nodes in each, and checks all the combinations of "/localnodes" requests with and without dc and rack options. This test also resolves a longstanding TODO that asked for such a multi-DC test for "/localnodes" to be written. Fixes #12147 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20915	2024-10-07 20:53:47 +03:00
Pavel Emelyanov	8bfbc563cc	test: Remove sstable factory from test_min_max_clustering_key() The helper makes sstables from env directly. Callers may not create the factor after that. Less code the better. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20983	2024-10-07 20:08:05 +03:00
Kefu Chai	a6ec6d32ab	auth: add "IWYU pragma: keep" to keep boost/regex_fwd.hpp clang-include-cleaner is not able to tell that the header provides the template parameter of `std::vector<std::pair<query_source, boost::regex>>`. and suggest us to remove this include. but it's wrong. so, in this change we apply the "pragma" to keep it. see https://github.com/include-what-you-use/include-what-you-use/blob/master/docs/IWYUPragmas.md for the explanations on what this pragma is for. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-07 20:08:05 +03:00
Kefu Chai	3d31835949	auth: include boost/regex_fwd.hpp in header since we only need the full definition of boost::regex in the .cc file, where we - define the constructor and destructor - and actually use the regex. there is no need to include boost/regex.hpp in the header, in order to keep the preprocessed header smaller. let's use a header only contains forward declarations in header, and include the full definition in the .cc file. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-07 20:08:05 +03:00
Piotr Smaron	ee56bbfe61	cql: sum of abs RFs diffs cannot exceed 1 in ALTER tablets KS Tablets load balancer is unable to process more than a single pending replica, thus ALTER tablets KS cannot accept an ALTER statement which would result in creating 2+ pending replicas, hence it has to validate if the sum of absoulte differences of RFs specified in the statement is not greter than 1.	2024-10-07 17:02:50 +02:00
Piotr Smaron	2aabe7f09c	cql: join new and old KS options in ALTER tablets KS A bug has been discovered while trying to ALTER tablets KS and specifying only 1 out of 2 DCs - the not specified DC's RF has been zeroed. This is because ALTER tablets KS updated the KS only with the RF-per-DC mapping specified in the ALTER tablets KS statement, so if a DC was ommitted, it was assigned a value of RF=0. This commit fixes that plus additionally passes all the KS options, not only the replication options, to the topology coordinator, where the KS update is performed. `initial_tablets` is a special case, which requires a special handling in the source code, as we cannot simply update old initial_tablet's settings with the new ones, because if only ` and TABLETS = {'enabled': true}` is specified in the ALTER tablets KS statement, we should not zero the `initial_tablets`, but rather keep the old value - this is tested by the `test_alter_preserves_tablets_if_initial_tablets_skipped` testcase. Other than that, the above mentioned testcase started to fail with these changes, and it appeared to be an issue with the test not waiting until ALTER is completed, and thus reading the old value, hence the test's body has been modified to wait for ALTER to complete before performing validation.	2024-10-07 17:02:45 +02:00
Avi Kivity	d12ba753e0	utils/unconst, mutation_partition: switch to ranges unconst is a small help that converts a const iterator to a non-const iterator with the help of the container. Currently it is using the boost iterator/range libraries. Convert it to <ranges> as part of an effort to standardize on a single range library. Its only user in mutation_partition is converted as well. Due to more iteroperability problems between <range> and boost, some calls to boost::adaptors::reversed have to be converted as well.	2024-10-07 17:30:12 +03:00
Avi Kivity	75f4ea1b68	utils: intrusive_btree: improve conformity with iterator requirements The <ranges> library checks that an iterator's operator++() returns a reference to the same type. intrusive_btree's iterator do not; instead they return some base type and rely on implicit conversion to the real iterator type. This causes interoperatibility problems with <range>. Fix by using the CRTP pattern to inform iterator_base about what type we really are, and cast to it. Enforce it with static_assert. Note we can't static_assert in class scope since it is checked too early and fails. Checking in function scope delays the check.	2024-10-07 17:26:01 +03:00
Piotr Smaron	6676e47371	cql: fix validation of ALTERing RFs in tablets KS The validation has been corrected with: 1. Checking if a DC specified in ALTER exists. 2. Removing `REPLICATION_STRATEGY_CLASS_KEY` key from a map of RFs that needs their RFs to be validated.	2024-10-07 16:02:01 +02:00
Piotr Smaron	93d61d7031	cql: harden `alter_keyspace_statement.cc::validate_rf_difference` This function assumed that strings passed as arguments will be of integer types, but that wasn't the case, and we missed that because this function didn't have any validation, so this change adds proper validation and error logging. Arguments passed to this function were forwarded from a call to `ks_prop_defs::get_replication_options`, which, among rf-per-dc mapping, returns also `class:replication_strategy` pair. Second pair's member has been casted into an `int` type and somehow the code was still running fine, but only extra testing added later discovered a bug in here.	2024-10-07 16:02:01 +02:00
Piotr Smaron	47acdc1f98	cql: validate RF change for new DCs in ALTER tablets KS ALTER tablets KS validated if RF is not changed by more than 1 for DCs that already had replicas, but not for DCs that didn't have them yet, so specifying an RF jump from 0 to 2 was possible when listing a new DC in ALTER tablets KS statement, which violated internal invariants of tablets load balancer. This PR fixes that bug and adds a multi-dc testcases to check if adding replicas to a new DC and removing replicas from a DC is honoring the RF change constraints. Refs: #20039	2024-10-07 16:02:01 +02:00
Piotr Smaron	9c5950533f	cql: extend test_alter_tablet_keyspace_rf Added cases to also test decreasing RF and setting the same RF. Also added extra explanatory comments.	2024-10-07 16:02:00 +02:00
Piotr Smaron	adf453af3f	cql: refactor test_tablets::test_alter_tablet_keyspace 1. Renamed the testcase to emphasize that it only focuses on testing changing RF - there are other tests that test ALTER tablets KS in general. 2. Fixed whitespaces according to PEP8	2024-10-07 16:02:00 +02:00
Piotr Smaron	042825247f	cql: remove unused helper function from test_tablets `change_default_rf` is not used anywhere, moreover it uses `replication_factor` tag, which is forbidden in ALTER tablets KS statement.	2024-10-07 16:02:00 +02:00
Nikos Dragazis	7a1ec3aa41	test: Test scrub/validate with SSTables from Cassandra All current unit tests for scrub in validate mode generate random SSTables on the fly. Add some more tests with frozen Cassandra SSTables from the source tree to verify compatibility with Cassandra. Use some of the existing 3.x Cassandra SSTables to test the valid case, and use the same schema to generate some corrupted SSTables for the invalid case. Overall, the new tests cover the following scenarios: * valid compressed/uncompressed * compressed/uncompressed with invalid checksums * compressed/uncompressed with invalid digest For the compressed SSTable with invalid checksums, a small chunk length was used (4KiB) to have more chunks with less disk space. For uncompressed SSTables the chunk length is not configurable. Finally, since the SSTables live in the source tree, the quarantine mechanism was disabled. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-07 15:21:38 +03:00
Nikos Dragazis	7090e2597f	compaction: Make quarantine optional for perform_sstable_scrub() Allow `perform_sstable_scrub()` to disable quarantine for invalid SSTables detected by scrub in validate mode. This is already supported by the lower-level function `scrub_sstables_validate_mode()` via the flag `quarantine_sstables` and is being used by sstable-scrub. Propagate the flag up to `perform_sstable_scrub()`. This will allow to test scrub/validate against read-only SSTables from the source tree. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-07 15:21:38 +03:00
Nikos Dragazis	5f2be2924e	test: Make random schema optional in scrub_test_framework The scrub_test_framework, which is the foundation for all scrub-related tests, always generates a random schema upon initialization and makes it available to the user. This is useful for running tests with ephemeral SSTables, but is redundant when the creation of the SSTable predates the test (e.g., it lives in the source tree). Turn scrub_test_framework into a template with a boolean parameter to optionally switch off the random schema generation. Also, add an overload for run() to support passing a ready-to-use SSTable instead of mutation fragments. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-07 15:21:38 +03:00
Nikos Dragazis	07ed0a48aa	test: Add tests for invalid digests In a previous patch we extended the validation path of the SSTable layer to validate the digests along with the checksums. Add two tests for compressed and uncompressed SSTables to test the validation API against SSTables with valid checksums but corrupted digests. Add two more tests to ensure that the absence of digest does not affect checksum validation. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-07 15:21:38 +03:00
Nikos Dragazis	39a74fb692	test: Merge scrub/validate tests for compressed and uncompressed cases Currently, every scrub/validate test is duplicated to cover both compressed and uncompressed SSTables. However, except for the compression type, the tests are identical. This leads to some code bloat. Introduce common functions parameterized by the compression type to reduce code duplication. Also, group together the compressed and uncompressed variants into one compression-agnostic test. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-07 15:21:38 +03:00
Nikos Dragazis	3a3783ee23	sstables: Verify digests on validation path Extend the validation path to perform digest checking on all SSTables. This is achieved by loading the digest component on demand and passing it to the underlying data sources only during validation. The data sources for compressed and uncompressed SSTables were modified in previous patches to support digest checking. Consider digest checking as part of the integrity checking mechanism (i.e., requires `integrity_check::yes`) to ensure it remains disabled for all reads happening outside of the validation path (i.e., `sstable::validate()`). This practically means that digest checking is enabled only for: * scrub in validate mode * sstable validate Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-07 15:21:09 +03:00
Avi Kivity	7dad248ac7	Merge 'Fix sstables registry mock' from Pavel Emelyanov There are two issues in it. First, listing the registry with a consumer callback passes wrong argument to the consumer. Second, the primary key of the registry is wrong. Both issues don't show up, because existing tests that use mock don't read from it, only write. Tests that read from registry are python tests that start scylla and thus use real registry. Closes scylladb/scylladb#20946 * github.com:scylladb/scylladb: test: Use corrcet key in sstables registry mock test: Pass entry status to mock registry consumer	2024-10-07 13:56:26 +03:00
Anna Stuchlik	a601845780	doc: remove outdated JMX references This commit removes references to JMX from the docs. Context: The JMX server has been dropped and removed from installation. The user can install it manually if needed, as documented with https://github.com/scylladb/scylladb/issues/18687. This commit removes the outdated information about JMX from other pages in the documentation, including the docs for nodetool, the list of ports, and the admin section. Also, the no longer relevant JMX information is removed from the Docker Hub docs. Fixes https://github.com/scylladb/scylladb/issues/18687 Fixes https://github.com/scylladb/scylladb/issues/19575 Closes scylladb/scylladb#20917	2024-10-07 13:55:15 +03:00
Nadav Har'El	987042be68	mv, test: reproduce missing validation for view name This patch adds reproducer tests (still failing) for issue #20755, which is about missing validation of materialized view names: 1. Unlike table and keyspace names which are limited to 48 characters, we forgot to limit view name length, and an excessively long name can cause Scylla to shut down :-( 2. Unlike table and keyspace names which only allow alphanumeric characters, view names are missing this check and can include any characters. 3. Luckily, even though we are missing the alphanumeric check, we at least don't allow "/" in view names (if we allowed them, it could allow users to write in any directory in the filesystem!). But when this happens, we get an internal error instead of the expected errors. The first test also fails on Cassandra (it doesn't crash it, but leaves the table in a strange state), but the other two pass. Refs #20755 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20761	2024-10-07 13:49:58 +03:00
Avi Kivity	73eeb6d274	Merge 'clang-format: adjustments to avoid unwanted refactors and better match the Seastar coding style' from Emil Maskovsky Some adjustments to the `.clang-format` options to better match the current code: * don't sort the include headers: causes large diffs especially in files with a lot of includes, and the `#include` ordering is not prescribed by the Seastar coding style * binpack the arguments in function declarations and calls: allow binpacking (as opposed to forcing each parameter on a separate line if they don't fit into the line length) * indented parameter continuation (as opposed to aligning to the open parenthesis) - aligning to the open parenthesis causes alignment issues especially with lambdas Fixes: scylladb/scylladb#20951 No backport: Not a product issue, just applies to master. Closes scylladb/scylladb#20968 * github.com:scylladb/scylladb: clang-format: argument and function packing clang-format: don't sort the include headers	2024-10-07 13:21:31 +03:00
Pavel Emelyanov	1870873538	test: Fix test_multiple_data_dirs The one was broken from the very beginning. It only checked that after creating a table, its directory is created in all datadirs. But it didn't check that after restart populating happens from the all. That's because all directories by 0th were always empty, so not-populating from them didn't skip any data. Fix it by moving all sstables from datadirs[0] to datadirs[1] before restart. With that update not-populating data from datadirs[1] will be noticed instantly. Fortunately, previous patches fixed that, so the test still passes. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	aa0c20a0e7	distributed_loader: Indentation fix after previous patch Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	792c0060c7	distributed_loader: Use correct datadir to collect local sstable Current code uses datadir it gets from table itself, which is the 0th element in the all-datadirs config. So populating local sstables happens several times from the same directory. Fix it by starting sstable directory with correct datadir -- the one obtained from the all-datadirs loop. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	bf654f45bd	distributed_loader: Move all-datadirs loop to local storage collecting It now happens in the outer loop, but it's not correct for S3 storage, which is thus asked to collect its data twice. Also it's broken for local storage as well, because the datadir argument is ignored. Next patch will fix it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	fe779ab1a2	distributed_loader: Collect table subdirs based on its storage options Collecting sstables for local storage and for S3 storage differs. First, the populator collects sstables for each datadir configured in scylla.yaml, but S3 storage doesn't care, so it's effectively asked to collect the same data twice. Second, S3 collector code uses sstable_directory simply because that class is used by reshape and reshard code, but in fact collecting of S3 sstable can be made much simpler (but that's for later). Having said that, split preparation of sstables population for local and S3 storage types. Indentation is deliberately left broken for local storage collecting mathod. That's because otherwise next patch will need move it back anyway. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	ee91cae5b9	distributed_loader: Indentation fix after previous patch Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	6ae486100c	distributed_loader: Squash loop of collect_subdir into one method This prepares the gound for the next patch. Indentation is left broken. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	4db2929afd	distributed_loader: Convert map of directories into a vector Knowledge of sstable state is no longer needed in the table_populator start/stop methods, so the map<state, directory> can be converted into vector<directory>. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	89e9231653	distributed_loader: Make start_subdir() method work with directory Similarly to populate_subdir() one, it also accepts state and gets directory out of it. Patch is the same way -- caller now passes it the reference to directory and doesn't care about the state (in fact, the start_subdir() doesn't care of the state either). While at it -- rename the method to reflect what it does. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	999ec88765	distributed_loader: Drop local reference variable Cleanup after previous patch. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	50024ef62a	distributed_loader: Split start_subdir() It does two things -- starts sstable_directory and prepares it. Split it accordingly. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	abdc0bb02d	distributed_loader: Remove allow-offstrategy argument This is to make populate_subdir() be self-contained in a way it uses passed sstable_directory and make caller not care about the state. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	3b583b9d9f	distributed_loader: Make populate() method work with directory The populate_subdir() accepts sstable_state argument and picks the corresponding sstable_directory object from the map. Patch it so that caller passes it the sstable_directory reference. For now it makes things more complicated, but next patches will simplify it back. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	051ac3a737	distributed_loader: Remove check for sstable_directory presense In the old days the set of sstable_directory-s used by populator could skip some of them. Now they are all present and the checks is always false. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	4752a504cd	distributed_loader: Out-line table_populator() methods To make further patching with less indentation level. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	617b0e3ce3	distributed_loader: Print storage options, not datadir Tables not necessarily have data in a directory, so it's more correct to show storage options in logs, not some directory path. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	31b2271f07	distributed_loader: Print prepared message When population throws, the catch block prepares a message to re-throw another exception and prints the same message into logs. Presumably the intent was to print the prepared message as well. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:04:23 +03:00
Pavel Emelyanov	87d392d071	sstable_directory: Add sstable_state argument ot one of constructors There's one constructor that became unused after `787ea4b1`. Modify it with the 'state' argument so that it could be used later. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 12:03:36 +03:00
Pavel Emelyanov	b56483ab67	sstable_directory: Add state() method The one will expose sstables state the directory works with. For convenience. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-07 11:23:50 +03:00
Kefu Chai	abda779a5b	compaction: return created sst without using a temporary variable simpler this way. `sst` does not help with the readability or performance, but let's drop it. simpler this way. also, remove the unused parameter. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20961	2024-10-07 10:56:25 +03:00
Pavel Emelyanov	8ccb4a1045	Merge 'db: remove unused includes ' from Kefu Chai these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. please note, since we have `using seastar::shared_ptr` in `seastarx.h`, this renders `#include <seastar/core/shared_ptr.hh>` unnecessary if we don't need the full definition of `seastar::shared_ptr`. so, in this change, all the unused includes are removed. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#20963 * github.com:scylladb/scylladb: .github: add db to iwyu's CLEANER_DIR db: remove unused includes	2024-10-07 10:55:48 +03:00
Kefu Chai	cd05f61607	api/storage_service: use ranges when handlging restore API this change is a follow up of `787ea4b1`, to modernize the code base. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20972	2024-10-07 10:54:37 +03:00
Avi Kivity	946bb870f3	utils: hashers: include <memory> hashers.hh uses std::unique_ptr, so include its header. Closes scylladb/scylladb#20974	2024-10-07 10:52:36 +03:00
Kefu Chai	c6bc5b2706	sstable_loader: Remove unused _snapshot_name from download_task_impl in `787ea4b1`, we introduced `_prefix` and `_sstables` member variables to `sstables_loader::download_task_impl`, replacing the functionality of `_snapshot_name`. However, we overlooked removing the now-obsolete `_snapshot_name` variable. this commit removes the unused `_snapshot_name` member variable to improve code cleanliness and prevent potential confusion. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20969	2024-10-07 10:43:13 +03:00
Benny Halevy	fa8fe62e90	test: cql-pytest: test_compaction: add test_compactionstats_after_major_compaction Test that compactionstats are empty, i.e. there are no required compactions following major compaction. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-07 10:24:06 +03:00
Benny Halevy	630b792bd0	test/cql-pytest: rename test_compaction{_tombstone_gc,} Prepare to add more tests related to compaction to this test suite. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-07 10:18:30 +03:00
Benny Halevy	284dbc51c3	time_window_compaction_strategy: estimated_pending_compactions: reestimate compactions rather than using cached value Currently, `estimated_pending_compactions` uses a precalculated value calculated by `update_estimated_compaction_by_tasks`, which, in turn, is called by `get_compaction_candidates`. That means that, if `estimated_pending_compactions` is called, e.g. right after major compaction, it will return an outdated value that was calculated prior to major compaction, and so, it is no longer relevant. Instead, just recalculate the value in `estimated_pending_compactions` and drop `update_estimated_compaction_by_tasks`. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-07 10:15:19 +03:00
Gleb Natapov	e642f0a86d	group: hold group0 shutdown gate during async operations Wait for all outstanding async work that uses group0 to complete before destroying group0 server. Fixes scylladb/scylladb#20701	2024-10-06 17:20:52 +03:00
Gleb Natapov	ba22493a69	group0: Stop group0 if node initialization fails Commit `af83c5e53e` moved aborting of group0 into the storage service drain function. But it is not called if node fails during initialization (if it failed to join cluster for instance). So lets abort on both paths (but only once).	2024-10-06 17:20:52 +03:00
Kefu Chai	960aa38cf3	utils/i_filter: include used header when compiling with clang-19 and the standard library from GCC-14.2, we have: ``` /usr/bin/cmake -E __run_co_compile --tidy="clang-tidy;--checks=-*,bugprone-use-after-move;--extra-arg-before=--driver-mode=g++" --source=/__w/scylladb/scylladb/utils/bloom_filter.cc -- /usr/bin/clang++ -DBOOST_REGEX_DYN_LINK -DBOOST_REGEX_NO_LIB -DFMT_SHARED -DSCYLLA_BUILD_MODE=release -DSEASTAR_API_LEVEL=7 -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SSTRING -DXXH_PRIVATE_API -I/__w/scylladb/scylladb -I/__w/scylladb/scylladb/seastar/include -I/__w/scylladb/scylladb/build/seastar/gen/include -I/__w/scylladb/scylladb/build/seastar/gen/src -ffunction-sections -fdata-sections -O3 -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-enum-constexpr-conversion -Wno-unused-parameter -ffile-prefix-map=/__w/scylladb/scylladb/build=. -march=wes Error: /__w/scylladb/scylladb/utils/bloom_filter.cc:81:1: error: unknown type name 'filter_ptr' [clang-diagnostic-error] 81 \| filter_ptr create_filter(int hash, large_bitset&& bitset, filter_format format) { \| ^ Error: /__w/scylladb/scylladb/utils/bloom_filter.cc:82:12: error: no viable conversion from returned value of type '__detail::__unique_ptr_t<murmur3_bloom_filter>' (aka 'unique_ptr<utils::filter::murmur3_bloom_filter>') to function return type 'int' [clang-diagnostic-error] 82 \| return std::make_unique<murmur3_bloom_filter>(hash, std::move(bitset), format); \| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Error: /__w/scylladb/scylladb/utils/bloom_filter.cc:85:1: error: unknown type name 'filter_ptr' [clang-diagnostic-error] 85 \| filter_ptr create_filter(int hash, int64_t num_elements, int buckets_per, filter_format format) { \| ^ Error: /__w/scylladb/scylladb/utils/bloom_filter.cc:86:12: error: no viable conversion from returned value of type '__detail::__unique_ptr_t<murmur3_bloom_filter>' (aka 'unique_ptr<utils::filter::murmur3_bloom_filter>') to function return type 'int' [clang-diagnostic-error] 86 \| return std::make_unique<murmur3_bloom_filter>(hash, large_bitset(get_bitset_size(num_elements, buckets_per)), format); \| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Error: /__w/scylladb/scylladb/utils/bloom_filter.hh:93:1: error: unknown type name 'filter_ptr' [clang-diagnostic-error] 93 \| filter_ptr create_filter(int hash, large_bitset&& bitset, filter_format format); \| ^ Error: /__w/scylladb/scylladb/utils/bloom_filter.hh:94:1: error: unknown type name 'filter_ptr' [clang-diagnostic-error] 94 \| filter_ptr create_filter(int hash, int64_t num_elements, int buckets_per, filter_format format); \| ^ Error: /__w/scylladb/scylladb/utils/i_filter.hh:17:25: error: no template named 'unique_ptr' in namespace 'std' [clang-diagnostic-error] 17 \| using filter_ptr = std::unique_ptr<i_filter>; \| ~~~~~^ Error: /__w/scylladb/scylladb/utils/i_filter.hh:54:12: error: unknown type name 'filter_ptr' [clang-diagnostic-error] 54 \| static filter_ptr get_filter(int64_t num_elements, double max_false_pos_prob, filter_format format); \| ^ 4 warnings and 8 errors generated. ``` apparently, the definition of `std::unique_ptr` is missing where it is used. so let's include `<memory>`, so that `i_filter.hh` is more self-contained. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20971	2024-10-06 14:20:41 +03:00
Michał Chojnowski	5884c9d2fc	utils/rjson.cc: correct a comment about assert() Commit `aa1270a00c` changed most uses of `assert` in the codebase to `SCYLLA_ASSERT`. But the comment fixed in this patch is talking specifically about `assert`, and shouldn't have been changed. It doesn't make sense after the change. Closes scylladb/scylladb#20967	2024-10-06 12:47:51 +03:00
Michał Chojnowski	882a3c60e4	utils/cached_file: reduce latency (and increase overhead) of partially-cached reads Currently, `cached_file::stream` (currently used only by index_reader, to read index pages), works as follows. Assume that the caller requested a read of the range [pos, pos + size). Then: - If the first page of the requested range is uncached, the entire [pos, pos + size) range is read from disk (even if some later pieces of it are cached), the resulting pages are added to the cache, and the read completes (most likely) from the cached pages. - If the first page of the read is cached, then the rest of the read is handled page-by-page, in a sequential loop, serving each page either from cache (if present) or from disk. For example, assume that pages 0, 1, 2, 3, 4 are requested. If exactly pages 1, 2 are cached, then `stream` will read the entire [0, 4] range from disk and insert the missing 0, 3, 4, and then it will continue serving the read from cache. If exactly pages 0 and 3 are cached, then it will serve 0 from cache, then it will read 1 from disk and insert it into cache, then it will read 2 from disk and insert it into cache, then it will serve 3 from cache, then it will read 4 from disk and insert it into cache. If exactly the first page is cached, a 128 kiB read turns into 31 I/O sequential read ops. This is weird, and doesn't look intended. In one case, we are reading even pages we already have, just to avoid fragmenting the read, and in the other case we are reading pages one-by-one (sequentially!) even if they are neighbours. I'm not sure if cached_file should minimize IOPS or byte throughput, but the current state is surely suboptimal. Even if its read strategy is somehow optimal, it should still at least coalesce contiguous reads and perform the non-contiguous reads in parallel. This patch leans into minimizing IOPS. After the patch, we serve as many front pages from the cache as we can, but when we see an uncached page, we read the entire remainder of the read from disk. As if we trimmed the read request by the longest cached prefix, and then performed the rest using the logic from before the patch. For example, if exactly pages 0 and 3 are cached, then we serve 0 from cache, then we read [1, 4] from disk and insert everything into cache. For partially-cached files, this will result in more bytes read from disk, but less IOPS. This might be a bad thing. But if so, then we should lean the other way in a more explicit and efficient way than we currently do. Closes scylladb/scylladb#20935	2024-10-04 17:39:38 +02:00
Emil Maskovsky	a11ede758e	clang-format: argument and function packing Changes to better match the Seastar code style and the current codebase. Allow parameter binpacking and continuation indenting. Refs: scylladb/scylladb#20951	2024-10-04 14:52:41 +02:00
Emil Maskovsky	b4f28b3e0e	clang-format: don't sort the include headers Sorting the include headers causes reordering of all headers and thus large diffs, especially in the files that include a lot of headers that have not been sorted before. This makes it harder to review the changes and to understand the history of the file. The Seastar code style doesn't prescribe any include headers ordering. Refs: scylladb/scylladb#20951	2024-10-04 14:51:54 +02:00
Kefu Chai	d72c8fc047	.github: add db to iwyu's CLEANER_DIR to avoid future violations of include-what-you-use. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-04 20:48:18 +08:00
Kefu Chai	ee36358a60	db: remove unused includes these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. please note, since we have `using seastar::shared_ptr` in `seastarx.h`, this renders `#include <seastar/core/shared_ptr.hh>` unnecessary if we don't need the full definition of `seastar::shared_ptr`. so, in this change, all the unused includes are removed. but there are some headers which are actually used, while still being identified by this tool. these includes are marked with "IWYU pragma: keep". Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-04 20:48:18 +08:00
Botond Dénes	af124993a4	Merge 'Do not remove objects from backup storage after restore' from Pavel Emelyanov The restore-from-s3 task uses load-and-stream internally which, in turn, unlinks loaded sstables on success. That's not what user expects when it restores from backup, objects should remain in bucket afterwards. Closes scylladb/scylladb#20947 * github.com:scylladb/scylladb: test: Add check that restored-from objects are not removed sstables_loader: Dont unlink sstables when restoring from S3 sstables_loader: Make primary_replica_only bool_class RAII field	2024-10-04 14:59:40 +03:00
Nikita Kurashkin	874cafefab	SStables: replace assertion with malformed_sstable_exception for invalid chunk_size This will allow to see underlying sstable file Fixes #20277 Closes scylladb/scylladb#20784	2024-10-04 14:48:35 +03:00
Pavel Emelyanov	6b480589fe	Merge 'treewide: accept list of sstables in "restore" API ' from Kefu Chai before this change, we enumerate the sstables tracked by the system.sstables table, and restore them when serving requests to "storage_service/restore" API. this works fine with "storage_service/backup" API. but this "restore" API cannot be used as a drop-in replacement of the rclone based API currently used by scylla-manager. in order to fill the gap, in this change: * add the "prefix" parameter for specifying the shared prefix of sstables * add the "sstables" parameter for specifying the list of TOC components of sstables * remove the "snapshot" parameter, as we don't encode the prefix on scylla's end anymore. * make the "table" parameter mandatory. Fixes https://github.com/scylladb/scylladb/issues/20461 ---- this change is a part of the efforts to bring the native backup/restore to scylla, no need to backprt. Closes scylladb/scylladb#20685 * github.com:scylladb/scylladb: treewide: accept list of sstables in "restore" API sstable: pass get_storage_option to sstable_directory::load_sstable() test/nodetool: add body parameter to `expected_request` tools/scylla-nodetool: enable nodetool to write HTTP body	2024-10-04 12:38:08 +03:00
Pavel Emelyanov	0f6e76f92f	api: Use captured compaction_manager in get_cm_stats() helper This is continuation of the previous patch that also need to touch the helper function argument. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-04 12:28:15 +03:00
Pavel Emelyanov	f99d8e07ae	api: Use captured compaction_manager in endpoints Instead of getting via ctx -> database chain. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-04 12:28:15 +03:00
Pavel Emelyanov	05b4a8e710	api: Add sharded<compaction_manager> argument to compaction_manager API reg/unreg To be used by next patch. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-04 12:28:15 +03:00
Pavel Emelyanov	58c4c21581	api: Move some endpoints from storage_service.cc to compaction_manager.cc Those setting and getting bandiwdth need compaction manager to work with and thus should sit next to other enpoints working with it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-04 11:36:07 +03:00
Pavel Emelyanov	43fe482204	api: Unset compaction_manager endpoints Similarly to other .cc files, compaction manager should have its endpoints unset. For now, no batch unsetting exists, so need to do it one-by-one. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-04 11:35:19 +03:00
Pavel Emelyanov	aa13be15b0	api: Use shorter registration method for compaction_manager function The register_api() helper does exatly what's needed here -- registers function and calls a method to set routes. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-04 11:34:33 +03:00
Botond Dénes	07094c3e44	Merge 'replica: Fix tombstone GC during tablet split preparation' from Raphael "Raph" Carvalho During split prepare phase, there will be more than 1 compaction group with overlapping token range for a given replica. Assume tablet 1 has sstable A containing deleted data, and sstable B containing a tombstone that shadows data in A. Then split starts: 1) sstable B is split first, and moved from main (unsplit) group to a split-ready group 2) now compaction runs in split-ready group before sstable A is split tombstone GC logic today only looks at underlying group, so compaction is step 2 will discard the deleted data in A, since it belongs to another group (the unsplit one), and so the tombstone can be purged incorrectly. To fix it, compaction will now work with all uncompacting sstables that belong to the same replica, since tombstone GC requires all sstables that possibly contain shadowed data to be available for correct decision to be made. Fixes https://github.com/scylladb/scylladb/issues/20044. Branches 6.0, 6.1 and 6.2 are vulnerable, so backport is needed. Closes scylladb/scylladb#20939 * github.com:scylladb/scylladb: replica: Fix tombstone GC during tablet split preparation service: Improve error handling for split	2024-10-04 10:29:42 +03:00
Nikos Dragazis	347f5ee166	sstables: Check if digest component exists Extend `read_digest()` to first check if the digest component exists before attempting to load it from disk. Make `validate_checksums()` throw an error if the component does not exist to preserve its current behavior. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-03 18:09:05 +03:00
Nikos Dragazis	7e738bcd2d	sstables: Add digest in the SSTable components SSTables store their digest in a Digest file. Add this in the list of SSTable components. In a follow-up patch we will use this component to enable digest checking in the validation path. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-03 18:09:05 +03:00
Nikos Dragazis	c893f06409	sstables: Add digest check in compressed data source Following the addition of digest check in the checksummed data source, add the same feature to the compressed data source as well. This ensures consistent behavior across any type of SSTable. This is added as an optional feature so that we can preserve the current behavior, that is verify only the per-chunk checksums during normal user reads. To ensure zero cost at runtime when disabled, we introduce the on/off switch as a template parameter. The digest calculation for compressed SSTables depends on the SSTable format, hence the new template argument for the checksum mode. This is consistent with the compressed data sink. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-03 18:09:01 +03:00
Nikos Dragazis	0df1c01759	sstables: Add digest check in checksummed data source The checksummed data source verifies the checksum of each chunk in the data files of uncompressed SSTables. This is being leveraged by scrub in validation mode. Extend the data source to check the digest (full checksum) as well. Unlike checksums, this is added as an optional feature so that SSTables without a digest can still be validated in a per-chunk basis. To enable this, the caller needs to set the template parameter `check_digest` to true, and provide the expected digest. The data source calculates the digest incrementally through multiple get() calls and compares against the expected digest after reading the whole file range. If there is a mismatch, it throws an exception. Checking the digest requires reading the whole data file. If this cannot be satisfied (e.g., due to partial read or skip()), the data source fails immediately. If the user has successfully read the whole file range, it can be safely assumed that the digest is valid. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-10-03 18:08:56 +03:00
Tomasz Grabiec	62f3d9e173	perf: perf_fast_forward: Add test case for querying missing rows	2024-10-03 16:26:41 +02:00
Tomasz Grabiec	4602ba90df	perf-fast-forward: Allow overriding promoted index block size For testing dense clustering index.	2024-10-03 16:26:41 +02:00
Tomasz Grabiec	1782456a52	perf-fast-forward: Test subsequent key reads from the middle in test_large_partition_select_few_rows	2024-10-03 16:26:41 +02:00
Tomasz Grabiec	751fa10de8	perf-fast-forward: Allow adding key offset in test_large_partition_select_few_rows	2024-10-03 16:26:41 +02:00
Tomasz Grabiec	10c6990e41	perf-fast-forward: Use single-partition reads in test_large_partition_select_few_rows It's a more realistic scenario than a full scan.	2024-10-03 16:26:28 +02:00
Tomasz Grabiec	753f6a61fd	sstables: bsearch_clustered_cursor: Add more tracing points	2024-10-03 16:24:18 +02:00
Botond Dénes	38088daa1f	scylla-gdb.py: drop compatibility code for EOL releases Any release < 6.0 or < 2023.1 is EOL and need not be supported by scylla-gdb.py anymore. Remove compatibility code for these releases. Closes scylladb/scylladb#20918	2024-10-03 15:42:08 +03:00
Avi Kivity	494561c4f3	cql3: expr: drop boost usage Replace boost usage with <ranges>, modernizing the code a little and reducing dependencies on a redundant library. Closes scylladb/scylladb#20919	2024-10-03 15:39:40 +03:00
Kefu Chai	7b82f3a375	test/lib: remove redundant fmt::to_string() in seastar::format() previously change, implementation was unnecessarily verbose and less efficient, as it created and immediately discarded temporary strings. remove unnecessary use of `fmt::to_string()` when arguments are already being formatted by `seastar::format()`. in this this change: - eliminates creation of temporary `std::string` instances - reduces memory allocations and copies - improves performance - simplifies the code Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20923	2024-10-03 15:36:55 +03:00
Tomasz Grabiec	95b864497a	sstables: reader: Log data file range	2024-10-03 14:16:05 +02:00
Tomasz Grabiec	41d3ae5e81	sstables: bsearch_clustered_cursor: Unify skip_info logging Now all exit paths which return skip_info will print it in the same way which makes for easier log parsing.	2024-10-03 14:16:05 +02:00
Tomasz Grabiec	1b82d5117a	sstables: bsearch_clustered_cursor: Narrow down range using "end" position of the block This is optimization. Example: block0: start=aaa, end=aaA block1: start=bbb, end=bbB block2: whatever Before the patch, advance_to("aAA") would skip to block0, and upper bound probe would skip to block1. This way, the reader would read the range of block0 from the data file. After the patch, "end" position is taken into account, so advance_to("aAA") will notice that block0 doesn't contain the position and will skip to block1. This is especially important for dense indexes, as it allows us to skip accessing data file if the search key is missing. It also solves the edge case problem related to the fact that single row reads are using a range which with positions which are not equal to the key, but are before(key) and after(key) for the lower bound and upper bound respectively. Before the patch, advance_to(before("bbb")) would skip to block0, before the position is before the block1's start. And upper bound probe for after("bbb") would point to block2. This way the read would scan block0 needlessly. After the patch, advance_to(before("bbb")) will skip to block1 because we notice based on "end" that block0 doesn't contain the position. This change also ensures that the start position of the upper bound entry of the after_key(pos), where pos is the last advance_to() position, is warm in cache. This is needed to optimize single-row reads with a dense index so that they always read exactly one promoted index block. For this to work, probe_upper_bound() for the after_key(row) always needs to find the upper bound block in cache.	2024-10-03 14:16:05 +02:00
Tomasz Grabiec	b03f23a09b	sstables: bsearch_clustered_cursor: Skip even to the first block It was unnecessary to emit a skip info for the first block since it follows immediately the partition start, but it is relevant to the optimization of avoiding data reads for missing keys. This optimization relies on the fact that lower bound position equals upper bound position. If the reader's key is before the first key in the partition and we don't arm the skip info for the first block, lower bound would be equal to the partition start, and upper bound would be equal to the first row's position, which are not equal.	2024-10-03 14:16:05 +02:00
Tomasz Grabiec	c905554121	test: sstables: sstable_3_x_test: Improve failure message	2024-10-03 14:16:05 +02:00
Tomasz Grabiec	7f077893ed	sstables: mx: writer: Never include partition_end marker in promoted index block width Currently, it may happen that the last promoted index block includes the partition_end marker. That's because we first write the partition end marker and then emit the unclosed block. This behavior matches Cassandra (checked in 3.x and 5.0.1). This is problematic for ruling out data file reads based on index. The width field is currently unused, but it will be used later where the width of the last block is used to compute the skip position past the last block for lookups which land after all keys in the partition. If width includes the marker then such a skip would land in the next partition, which is incorrect, as the reader context expects a cell element. Even if that was recognized, it's wrong - if this is not a single partition read (so upper bound is not at the next partition too), then we would read from the wrong (next) partition. We want to be able to make such skips in order to avoid unnecessary data file IO for reads of missing rows. Currently, we would always read the last block even if the key is past its "end" position. Another way to solve this would be to propagate the "past the last block" condition from the index cursor to the reader and let it deal with it, but the logic for that would be complicated. With this fix, there is no special logic required.	2024-10-03 14:09:57 +02:00
Pavel Emelyanov	4465bd9e5e	test: Add check that restored-from objects are not removed Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-03 13:37:04 +03:00
Botond Dénes	3ebb124eb2	repair/row_level: remove reader timeout This timeout was added to catch reader related deadlocks. We have not seen such deadlocks for a long time, but we did see false-timeouts caused by this, see explanation below. Since the cost now outweight the benefit, remove the timeout altogether. The false timeout happens during mixed-shard repair. The `reader_permit::set_timeout()` call is called on the top-level permit which repair has a handle on. In the case of the mixed-shard repair, this belongs to the multishard reader. Calling set_timeout() on the multishard reader has no effect on the actual shard readers, except in one case: when the shard reader is created, it inherits the multishard reader's current timeout. As the shard reader can be alive for a long time, this timeout is not refreshed and ultimately causes a timeout and fails the repair. Refs: #18269 Closes scylladb/scylladb#20703	2024-10-03 11:26:29 +02:00
Kamil Braun	e67016540c	Merge 'Node replace and remove operations: Add deprecate IP addresses usage warning.' from Sergey Zolotukhin - As part of deprecation of IP address usage, warning messages were added when IP addresses specified in the `ignore-dead-nodes` and `--ignore-dead-nodes-for-replace` options for scylla and nodetool. - Slight optimizations for `utils::split_comma_separated_list`, ` host_id_or_endpoint lists` and `storage_service` remove node operations, replacing `std::list` usage with `std::vector`. Fixes scylladb/scylladb#19218 Backport: 6.2 as it's not yet released. Closes scylladb/scylladb#20756 * github.com:scylladb/scylladb: config: Add a warning about use of IP address for join topology and replace operations. nodetool: Add IP address usage warning for 'ignore-dead-nodes'. tests: Fix incorrect UUIDs in test_nodeops utils: Optimizations for utils::split_comma_separated_list and usage of host_id_or_endpoint lists	2024-10-03 11:08:28 +02:00
Kamil Braun	d2233d4400	Merge 'test: update cql/ tests to work with tablets enabled by default' from Konstantin Osipov Explicitly disable tablets for features which still dont' work with tablets: cdc, lwt, coutners. Closes scylladb/scylladb#20858 * github.com:scylladb/scylladb: test: make cdc tests pass with tablets on by default test: make cql/counters* pass with and without tablets test: make cql/lwt_* pass with and without tablets test: rename cql/list_test to cql/lwt_list_test	2024-10-03 10:53:17 +02:00
Kefu Chai	f9091066b7	treewide: replace boost::irange with std::views::iota where possible when building scylla with the standard library from GCC-14.2, shipped by fedora 41, we have following build failure: ``` /home/kefu/.local/bin/clang++ -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -isystem /home/kefu/dev/scylladb/abseil -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb/build=. -march=x86-64-v3 -mpclmul -Xclang -fexperimental-assignment-tracking=disabled -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -MD -MT CMakeFiles/scylla-main.dir/Debug/init.cc.o -MF CMakeFiles/scylla-main.dir/Debug/init.cc.o.d -o CMakeFiles/scylla-main.dir/Debug/init.cc.o -c /home/kefu/dev/scylladb/init.cc In file included from /home/kefu/dev/scylladb/init.cc:12: In file included from /home/kefu/dev/scylladb/db/config.hh:20: In file included from /home/kefu/dev/scylladb/locator/abstract_replication_strategy.hh:26: /home/kefu/dev/scylladb/locator/tablets.hh:410:30: error: unexpected type name 'size_t': expected expression 410 \| return boost::irange<size_t>(0, tablet_count()) \| boost::adaptors::transformed([] (size_t i) { \| ^ /home/kefu/dev/scylladb/locator/tablets.hh:410:23: error: no member named 'irange' in namespace 'boost' 410 \| return boost::irange<size_t>(0, tablet_count()) \| boost::adaptors::transformed([] (size_t i) { \| ~~~~~~~^ /home/kefu/dev/scylladb/locator/tablets.hh:410:38: error: left operand of comma operator has no effect [-Werror,-Wunused-value] 410 \| return boost::irange<size_t>(0, tablet_count()) \| boost::adaptors::transformed([] (size_t i) { \| ^ 3 errors generated. [16/782] Building CXX object CMakeFiles/scylla-main.dir/Debug/keys.cc.o [17/782] Building CXX object CMakeFiles/scylla-main.dir/Debug/counters.cc.o [18/782] Building CXX object CMakeFiles/scylla-main.dir/Debug/partition_slice_builder.cc.o [19/782] Building CXX object CMakeFiles/scylla-main.dir/Debug/mutation_query.cc.o FAILED: CMakeFiles/scylla-main.dir/Debug/mutation_query.cc.o /home/kefu/.local/bin/clang++ -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -isystem /home/kefu/dev/scylladb/abseil -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb/build=. -march=x86-64-v3 -mpclmul -Xclang -fexperimental-assignment-tracking=disabled -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -MD -MT CMakeFiles/scylla-main.dir/Debug/mutation_query.cc.o -MF CMakeFiles/scylla-main.dir/Debug/mutation_query.cc.o.d -o CMakeFiles/scylla-main.dir/Debug/mutation_query.cc.o -c /home/kefu/dev/scylladb/mutation_query.cc In file included from /home/kefu/dev/scylladb/mutation_query.cc:12: In file included from /home/kefu/dev/scylladb/schema/schema_registry.hh:17: In file included from /home/kefu/dev/scylladb/replica/database.hh:11: In file included from /home/kefu/dev/scylladb/locator/abstract_replication_strategy.hh:26: /home/kefu/dev/scylladb/locator/tablets.hh:410:30: error: unexpected type name 'size_t': expected expression 410 \| return boost::irange<size_t>(0, tablet_count()) \| boost::adaptors::transformed([] (size_t i) { \| ^ /home/kefu/dev/scylladb/locator/tablets.hh:410:23: error: no member named 'irange' in namespace 'boost' 410 \| return boost::irange<size_t>(0, tablet_count()) \| boost::adaptors::transformed([] (size_t i) { \| ~~~~~~~^ /home/kefu/dev/scylladb/locator/tablets.hh:410:38: error: left operand of comma operator has no effect [-Werror,-Wunused-value] 410 \| return boost::irange<size_t>(0, tablet_count()) \| boost::adaptors::transformed([] (size_t i) { \| ^ In file included from /home/kefu/dev/scylladb/mutation_query.cc:12: In file included from /home/kefu/dev/scylladb/schema/schema_registry.hh:17: In file included from /home/kefu/dev/scylladb/replica/database.hh:37: In file included from /home/kefu/dev/scylladb/db/snapshot-ctl.hh:20: /home/kefu/dev/scylladb/tasks/task_manager.hh:403:54: error: no member named 'irange' in namespace 'boost' 403 \| co_await coroutine::parallel_for_each(boost::irange(0u, smp::count), [&tm, id, &res, &func] (unsigned shard) -> future<> { \| ~~~~~~~^ 4 errors generated. ``` so let's take the opportunity to switch from `boost::irange` to `std::views::iota`. in this change, we: - switch from boost::irange to std::views::iota for better standard library compatibility - retain boost::irange where step parameter is used, as std::views::iota doesn't support it - this change partially modernizes our range usage while maintaining - existing functionality Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20924	2024-10-03 10:33:33 +03:00
Pavel Emelyanov	7389f4275d	sstables_loader: Dont unlink sstables when restoring from S3 When load_and_stream() completes, all sstables that were loaded (and streamed) are unlinked. This is wrong for the restore-from-s3 task, as removing objects from backup storage is not what user expects. Fix it by adding a boolean to streamer class, and set it to false (well, bool_class<>::no) for restore task. fixes: #20938 Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-03 10:15:22 +03:00
Pavel Emelyanov	7eb48358e9	sstables_loader: Make primary_replica_only bool_class RAII field This boolean is currently passed all the way around as pure bool argument. And it's only needed in a single get_endpoints() method that calculates the target endpoints. This patch places this bool on class streamer, so that the call chain arguments are not polluted, and converts it to bool_class. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-03 10:13:37 +03:00
Pavel Emelyanov	1da1d131b2	test: Use corrcet key in sstables registry mock The "real" registry defines its primary key as (location, generation) pair, where location is the partition key and generation is clustering key. The registry mock uses only location part as primary key, while it must use both. The buggy mock works simply because the listing API is in fact not used by unit tests. Those tests that do need it are python tests that start scylla and thus implicitly use real registry. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-03 10:07:11 +03:00
Pavel Emelyanov	a503e2ab10	test: Pass entry status to mock registry consumer When sstables registry is listed, the passed consumer accepts entry status as its first argument, not its location (location is passed as a search key) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-10-03 10:06:28 +03:00
Kefu Chai	c7eafc4dc1	auth: capture boost::regex_error not std::regex_error in `a3db5401`, we introduced the TLS certi authenticator, which is configured using `auth_certificate_role_queries` option . the value of this option contains a regular expression. so there are chances the regular expression is malformatted. in that case, when converting its value presenting the regular expression to an instance of `boost::regex`, Boost.Regex throws a `boost::regex_error` exception, not `std::regex_error`. since we decided to use Boost.Regex, let's catch `boost::regex_error`. Refs `a3db5401` Fixes scylladb/scylladb#20941 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20942	2024-10-03 09:57:15 +03:00
Piotr Dulikowski	6778001313	Merge 'cql3: Make creating MV respect ID option' from Dawid Mędrek Before these changes, we could create a materialized view specifying its ID, but the option was ignored. This commit makes Scylla respect the option. Now specifying the ID results in the MV being created with that specific ID. This way, Scylla's behavior is consistent with Cassandra's. Because Cassandra doesn't mention the option in its user documentation, we don't update it either in case the semantics of it changes in the future -- we want to have an open door for any modifications. Note that Cassandra returns a server error if the provided ID is already in use, both in the case of regular tables and MVs. That's most likely a bug. Instead of following that behavior, we stay consistent with the current semantics of creating a regular table in Scylla: if the provided ID is already used, return an InvalidRequest. The last thing worth pointing out is Cassandra handles `WITH ID = null` as a special case; normally, specifying an invalid ID results in a ConfigurationException, but a null is treated as a syntax error. As in the previous paragraph, we stay consistent with the semantics of regular tables and all invalid IDs, null included, lead to a ConfigurationException. We also add a few short tests verifying that the implementation works as intended. Fixes scylladb/scylladb#20616 Backport not needed: the semantics of the option was never documented in either Cassandra, or Scylla. Closes scylladb/scylladb#20773 * github.com:scylladb/scylladb: test/cql-pytest: Get rid of unnecessary processing describe statements cql3: Make creating MV respect ID option	2024-10-03 08:31:07 +02:00
Dawid Mędrek	1f1b201fd8	cql3/functions/user_function: Use fmt to format create statement We replace `std::ostringstream` with views and formatting using fmt to improve readability of the code.	2024-10-02 19:17:35 +02:00
Ferenc Szili	cdf775d3cc	test: test tombstone GC disabled on pending replica This tests if tombstone GC is disabled on pending replicas	2024-10-02 16:37:57 +02:00
Ferenc Szili	ba6707506d	tablet_storage_group_manager: update tombstone_gc_enabled in compaction group In order to avoid cases during tablet migrations where we garbage collect tombstones before the data it shadows arrives, we will disable tombstone GC on pending replicas. To achieve this we added a tombston_gc_enabled flag to compaction_group. This flag is updated from updte_effective_repliction_map method of the tablet_storage_group_manager class.	2024-10-02 16:31:33 +02:00
Raphael S. Carvalho	93815e0649	replica: Fix tombstone GC during tablet split preparation During split prepare phase, there will be more than 1 compaction group with overlapping token range for a given replica. Assume tablet 1 has sstable A containing deleted data, and sstable B containing a tombstone that shadows data in A. Then split starts: 1) sstable B is split first, and moved from main (unsplit) group to a split-ready group 2) now compaction runs in split-ready group before sstable A is split tombstone GC logic today only looks at underlying group, so compaction is step 2 will discard the deleted data in A, since it belongs to another group (the unsplit one), and so the tombstone can be purged incorrectly. To fix it, compaction will now work with all uncompacting sstables that belong to the same replica, since tombstone GC requires all sstables that possibly contain shadowed data to be available for correct decision to be made. Fixes #20044. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-10-02 11:26:13 -03:00
Ferenc Szili	e472844a78	database::table: add tombstone_gc_enabled(locator::tablet_id) This change adds the flag tombstone_gc_enabled to compaction_group. The value of this flag will be set in tablet_storage_group_manager::update_effective_replication_map().	2024-10-02 16:24:45 +02:00
Raphael S. Carvalho	bcd358595f	service: Improve error handling for split Retry wasn't really happening since the loop was broken and sleep part was skipped on error. Also, we were treating abort of split during shutdown as if it were an actual error and that confused longevity tests that parse for logs with error level. The fix is about demoting the level of logs when we know the exception comes from shutdown. Fixes #20890.	2024-10-02 11:23:44 -03:00
Konstantin Osipov	1d1777b13a	test: make cdc tests pass with tablets on by default CDC is not supported with tablets, explicitly disable tablets in CDC keyspace definition.	2024-10-02 06:37:14 -04:00
Konstantin Osipov	0e3dbec277	test: make cql/counters* pass with and without tablets Counters are not supported with tablets, make sure the test works in any ScyllaDB configuration.	2024-10-02 06:37:14 -04:00
Konstantin Osipov	92aca17bc5	test: make cql/lwt_* pass with and without tablets Lightweight transactions don't support tablets, so let's explicitly disable tablets in LWT tests.	2024-10-02 06:37:14 -04:00
Konstantin Osipov	7c64fc0c4f	test: rename cql/list_test to cql/lwt_list_test This test is actually testing lists with LWT, so should have the corresponding name. Going forward we'll patch CQL LWT tests for tablets, so let's group them together.	2024-10-02 06:37:14 -04:00
Sergey Zolotukhin	6398b7548c	config: Add a warning about use of IP address for join topology and replace operations. When the '--ignore-dead-nodes-for-replace' config option contains IP addresses, a warning will be logged, notifying the user that using IP addresses with this option is deprecated and will no longer be supported in the next release. Fixes scylladb/scylladb#19218	2024-10-02 11:56:59 +02:00
Sergey Zolotukhin	9c692438e9	nodetool: Add IP address usage warning for 'ignore-dead-nodes'. Since we are deprecating the use of IP addresses, a warning message will be printed if 'nodetool removenode --ignore-dead-nodes' is used with IP addresses.	2024-10-02 11:56:59 +02:00
Sergey Zolotukhin	a871321ecf	tests: Fix incorrect UUIDs in test_nodeops It was found that the UUIDs used in test_nodeops were invalid. This update replaces those UUIDs with newly generated random UUIDs.	2024-10-02 11:56:59 +02:00
Sergey Zolotukhin	3b9033423d	utils: Optimizations for utils::split_comma_separated_list and usage of host_id_or_endpoint lists - utils::split_comma_separated_list now accepts a reference to sstring instead of a copy to avoid extra memory allocations. Additionally, the results of trimming are moved to the resulting vector instead of being copied. - service/storage_service removenode, raft_removenode, find_raft_nodes_from_hoeps, parse_node_list and api/storage_service::set_storage_service were changed to use std::vector<host_id_or_endpoint> instead of std::list<host_id_or_endpoint> as std::vector is a more cache-friendly structure, resulting in better performance.	2024-10-02 11:56:59 +02:00
Dawid Mędrek	7a7a1e3558	treewide: Prefer bytes_fwd.hh over bytes.hh CI started reporting warnings about including `bytes.hh` in several files. The reason is they actually only use code introduced in `bytes_fwd.hh` (which is also included by `bytes.hh`). Clang-include-cleaner suggests that we get rid of that indirection and only include `bytes_fwd.hh`. That's what happens in this commit. We include `bytes.hh` in `exceptions/exceptions.cc` because it relies on the formatting utilities declared and defined in `bytes.hh`. Closes scylladb/scylladb#20842	2024-10-02 07:29:30 +02:00
Dawid Mędrek	de88c150f6	test/cql-pytest: Get rid of unnecessary processing describe statements As part of scylladb/scylladb@d42f160, we added a test verifying that restoring the schema works as intended. Unfortunately, because of scylladb/scylladb#20616, we had to manually process the results of `DESCRIBE SCHEMA` to exclude the ID parameter and be able to compare restore statements corresponding to the same view. Now that materialized views respect the ID parameter, we can get rid of that logic.	2024-10-01 22:04:05 +02:00
Dawid Mędrek	552c752005	cql3: Make creating MV respect ID option Before these changes, we could create a materialized view specifying its ID, but the option was ignored. This commit makes Scylla respect the option. Now specifying the ID results in the MV being created with that specific ID. This way, Scylla's behavior is consistent with Cassandra's. Because Cassandra doesn't mention the option in its user documentation, we don't update it either in case the semantics of it changes in the future -- we want to have an open door for any modifications. Note that Cassandra returns a server error if the provided ID is already in use, both in the case of regular tables and MVs. That's most likely a bug. Instead of following that behavior, we stay consistent with the current semantics of creating a regular table in Scylla: if the provided ID is already used, return an InvalidRequest. The last thing worth pointing out is Cassandra handles `WITH ID = null` as a special case; normally, specifying an invalid ID results in a ConfigurationException, but a null is treated as a syntax error. As in the previous paragraph, we stay consistent with the semantics of regular tables and all invalid IDs, null included, lead to a ConfigurationException. We also add a few short tests verifying that the implementation works as intended.	2024-10-01 22:03:58 +02:00
Kefu Chai	9b5eab0dde	test/lib: include <fmt/std.h> for formatting std::optional before this change, when compiling with fmtlib v11.0.2 and clang v19.1.0, the compiler fails like: ``` /usr/bin/clang++ -DBOOST_REGEX_DYN_LINK -DBOOST_REGEX_NO_LIB -DBOOST_UNIT_TEST_FRAMEWORK_DYN_LINK -DBOOST_UNIT_TEST_FRAMEWORK_NO_LIB -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -I/home/kefu/dev/scylladb/build -isystem /home/kefu/dev/scylladb/abseil -isystem /home/kefu/dev/scylladb/build/rust -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-enum-constexpr-conversion -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb/build=. -march=x86-64-v3 -mpclmul -Xclang -fexperimental-assignment-tracking=disabled -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -MD -MT test/lib/CMakeFiles/test-lib.dir/Debug/cql_assertions.cc.o -MF test/lib/CMakeFiles/test-lib.dir/Debug/cql_assertions.cc.o.d -o test/lib/CMakeFiles/test-lib.dir/Debug/cql_assertions.cc.o -c /home/kefu/dev/scylladb/test/lib/cql_assertions.cc In file included from /home/kefu/dev/scylladb/test/lib/cql_assertions.cc:12: In file included from /usr/include/fmt/ranges.h:20: In file included from /usr/include/fmt/format.h:41: /usr/include/fmt/base.h:2673:45: error: implicit instantiation of undefined template 'fmt::detail::type_is_unformattable_for<std::vector<std::optional<seastar::basic_sstring<signed char, unsigned int, 31, false>>>, char>' 2673 \| type_is_unformattable_for<T, char_type> _; \| ^ /usr/include/fmt/base.h:2735:23: note: in instantiation of function template specialization 'fmt::detail::parse_format_specs<std::vector<std::optional<seastar::basic_sstring<signed char, unsigned int, 31, false>>>, fmt::detail::compile_parse_context<char>>' requested here 2735 \| parse_funcs_{&parse_format_specs<Args, parse_context_type>...} {} \| ^ /usr/include/fmt/base.h:2884:47: note: in instantiation of member function 'fmt::detail::format_string_checker<char, int, std::vector<std::optional<seastar::basic_sstring<signed char, unsigned int, 31, false>>>, std::vector<std::optional<managed_bytes>>>::format_string_checker' requested here 2884 \| detail::parse_format_string<true>(str_, checker(s)); \| ^ /home/kefu/dev/scylladb/test/lib/cql_assertions.cc:132:34: note: in instantiation of function template specialization 'fmt::basic_format_string<char, int &, std::vector<std::optional<seastar::basic_sstring<signed char, unsigned int, 31, false>>> &, const std::vector<std::optional<managed_bytes>> &>::basic_format_string<char[35], 0>' requested here 132 \| fail(seastar::format("row {} differs, expected {} got {}", row_nr, row, actual)); \| ^ /usr/include/fmt/base.h:1616:8: note: template is declared here 1616 \| struct type_is_unformattable_for; \| ^ /home/kefu/dev/scylladb/test/lib/cql_assertions.cc:132:34: error: call to consteval function 'fmt::basic_format_string<char, int &, std::vector<std::optional<seastar::basic_sstring<signed char, unsigned int, 31, false>>> &, const std::vector<std::optional<managed_bytes>> &>::basic_format_string<char[35], 0>' is not a constant expression 132 \| fail(seastar::format("row {} differs, expected {} got {}", row_nr, row, actual)); \| ^ ``` because the formatter for `std::optional<>` is defined in fmt/std.h. so, in this change, we include the used header. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20922	2024-10-01 22:32:16 +03:00
Tomasz Grabiec	a29501ed67	sstables: Reduce amount of I/O for clustering-key-bounded reads from large partitions Single-row reads from large partition issue 64 KiB reads to the data file, which is equal to the default span of the promoted index block in the data file. If users would want to reduce selectivity of the index to speed up single-row reads, this won't be effective. The reason is that the reader uses promoted index to look up the start position in the data file of the read, but end position will in practice extend to the next partition, and amount of I/O will be determined by the underlying file input stream implementation and its read-ahead heuristics. By default, that results in at least 2 IOs 32KB each. There is already infrastructure to lookup end position based on upper bound of the read, but it's not effective becasue it's a non-populating lookup and the upper bound cursor has its own private cached_promoted_index, which is cold when positions are computed. It's non-populating on purpose, to avoid extra index file IO to read upper bound. In case upper bound is far-enough from the lower bound, this will only increase the cost of the read. The solution employed here is to warm up the lower bound cursor's cache before positions are computed, and use that cursor for non-populating lookup of the upper bound. We use the lower bound cursor and the slice's lower bound so that we read the same blocks as later lower-bound slicing would, so that we don't incur extra IO for cases where looking up upper bound is not worth it, that is when upper bound is far from the lower bound. If upper bound is near lower bound, then warming up using lower bound will populate cached_promoted_index with blocks which will allow us to locate the upper bound block accurately. This is especially important for single-row reads, where the bounds are around the same key. In this case we want to read the data file range which belongs to a single promoted index block. It doesn't matter that the upper bound is not exactly the same. They both will likely lie in the same block, and if not, binary search will bring adjacent blocks into cache. Even if upper bound is not near, the binary search will populate the cache with blocks which can be used to narrow down the data file range somewhat. Fixes #10030. The change was tested with perf-fast-forward. I populated the data set with `column_index_size_in_kb` set to 1 scylla perf-fast-forward --populate --run-tests=large-partition-slicing --column-index-size-in-kb=1 Test run: build/release/scylla perf-fast-forward --run-tests=large-partition-select-few-rows -c1 --keep-cache-across-test-cases --test-case-duration=0 This test reads two rows from the middle of a large partition (1M rows), of subsequent keys. The first read will miss in the index file page cache, the second read will hit. Notice that before the change, the second read issued 2 aio requests worth of 64KiB in total. After the change, the second read issued 1 aio worth of 2 KiB. That's because promoted index block is larger than 1 KiB. I verified using logging that the data file range matches a single promoted index block. Also, the first read which misses in cache is still faster after the change. Before: running: large-partition-select-few-rows on dataset large-part-ds1 Testing selecting few rows from a large partition: stride rows time (s) iterations frags frag/s mad f/s max f/s min f/s avg aio aio (KiB) blocked dropped idx hit idx miss idx blk c hit c miss c blk allocs tasks insns/f cpu 500000 1 0.009802 1 1 102 0 102 102 21.0 21 196 2 1 0 1 1 0 0 0 568 269 4716050 53.4% 500001 1 0.000321 1 1 3113 0 3113 3113 2.0 2 64 1 0 1 0 0 0 0 0 116 26 555110 45.0% After: running: large-partition-select-few-rows on dataset large-part-ds1 Testing selecting few rows from a large partition: stride rows time (s) iterations frags frag/s mad f/s max f/s min f/s avg aio aio (KiB) blocked dropped idx hit idx miss idx blk c hit c miss c blk allocs tasks insns/f cpu 500000 1 0.009609 1 1 104 0 104 104 20.0 20 137 2 1 0 1 1 0 0 0 561 268 4633407 43.1% 500001 1 0.000217 1 1 4602 0 4602 4602 1.0 1 2 1 0 1 0 0 0 0 0 110 26 313882 64.1% (cherry picked from commit dfb339376aff1ed961b26c4759b1604f7df35e54)	2024-10-01 18:40:34 +02:00
Tomasz Grabiec	41be5d1daf	sstables: clustered_cursor: Track current block Will be needed by the reader to jump to the current block even if we already advanced to it before, when setting up the reader context. We want to advance to lower bound earlier, before the praser skips to the lower bound. We want that in order to set input stream data file range based on index. If we didn't have access to the current block and used the result from advance_to(), the parser will think we're already in the block which has lower_bound when it attempts to skip, and will not skip, falling back to scanning.	2024-10-01 18:40:34 +02:00
Kefu Chai	787ea4b1d4	treewide: accept list of sstables in "restore" API before this change, we enumerate the sstables tracked by the system.sstables table, and restore them when serving requests to "storage_service/restore" API. this works fine with "storage_service/backup" API. but this "restore" API cannot be used as a drop-in replacement of the rclone based API currently used by scylla-manager. in order to fill the gap, in this change: * add the "prefix" parameter for specifying the shared prefix of sstables * add the "sstables" parameter for specifying the list of TOC components of sstables * remove the "snapshot" parameter, as we don't encode the prefix on scylla's end anymore. * make the "table" parameter mandatory. Fixes scylladb/scylladb#20461 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-01 23:24:56 +08:00
Kefu Chai	17181c2eca	sstable: pass get_storage_option to sstable_directory::load_sstable() before this change, we always pass `sstable_directory::_storage_opts` to `_manager.make_sstable()` in `sstable_directory::load_sstable()`. but when loading from object storage, we need to customize the storage_options on a per-sstable basis. the way to address this is to allow the caller of `sstable_directory::process_descriptor()` to pass a functor which return the `storage_options` to be used when creating the sstable. so, in this change, we update - sstable_directory::load_sstable() - sstable_directory::process_descriptor() so that they accept another parameter to create the storage_options. in the next commit we will pass a different functor for customizing the storage_options on a per-sstable basis when loading sstables. Refs scylladb/scylladb#20461 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-01 23:24:56 +08:00
Kefu Chai	283697e316	test/nodetool: add body parameter to `expected_request` before this change, `expected_request` only includes query strings for the parameters of requests. but we will add an API ("storage_service/restore") which accepts its parameters in HTTP body as well. in this change, we add an optional `body` member to `expected_request`, so that we can mock the APIs which pass the parameters with the HTTP body. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-01 23:24:56 +08:00
Kefu Chai	3c19cc9aec	tools/scylla-nodetool: enable nodetool to write HTTP body before this change, we always send the parameters with query strings, but we will add an API ("storage_service/restore") which accepts its parameters in HTTP body as well. in this change, we add an optional parameter to `do_request()` and `post()`, so that we can send HTTP body when using "POST" method in nodetool implementation. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-01 23:24:56 +08:00
Pavel Emelyanov	7f71371de1	distributed_loader: Get token metadata from e.r.m., not database Though database can be used to get relevant token metadata, it's better not to use one service (database) as a proxy to get another one (token metadata). In case of tokens, there's effective replication map at hand, which is a more correct source of such topology information. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20894	2024-10-01 14:59:35 +03:00
Anna Stuchlik	7eb1dc2ae5	doc: document the option to run ScyllaDB in Docker on macOS This commit adds a description of a workaround to create a multi-node ScyllaDB cluster with Docker on macOS. Refs https://github.com/scylladb/scylladb/issues/16806 See https://forum.scylladb.com/t/running-3-node-scylladb-in-docker/1057/4 Closes scylladb/scylladb#20857	2024-10-01 14:58:58 +03:00
Botond Dénes	6535283881	.github/CODEOWNERS: add code owners for tools/* Closes scylladb/scylladb#20702	2024-10-01 14:52:26 +03:00
Yaron Kaikov	ab964bcd5a	[script/pull_github_pr.sh] Check Gating status before merging Maintainers use scripts/pull_github_pr.sh from scylladb.git when merging PRs and before pushing to the next. We want to prevent merges from piling up on top of unstable builds. This change will check Gating's current status and notify the maintainers Related to scylladb/scylla-pkg#3644 Closes scylladb/scylladb#20742	2024-10-01 14:46:29 +03:00
Anna Stuchlik	a97db03448	doc: add metric updates from 6.1 to 6.2 This commit specifies metrics that are new in version 6.2 compared to 6.1, as specified in https://github.com/scylladb/scylladb/issues/20176. Fixes https://github.com/scylladb/scylladb/issues/20176 Closes scylladb/scylladb#20896	2024-10-01 14:41:37 +03:00
muthu90tech	1204d54c5c	transport: Dont bypass seastar API when making syscalls The transport/controller.cc bypasses seastar API when making a few syscalls, this PR will use the right seastar API to make the syscall and libc calls this PR relies on few new APIs introduced in seastar commit : cd7f3b8e8850cd80a4f6899cedc726e576c51abe Closes scylladb/scylladb#17443 Closes scylladb/scylladb#19565	2024-10-01 14:29:24 +03:00
Benny Halevy	5a0f3889e0	treewide: use std::ranges sort functions rather than boost Using the standard library is preffered over boost. In cql3/expr/expression.cc to_sorted_vector got more of a face-list and was modernized to use also std::unique and while at it, to move its input range in the uniquely sorted result vector. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-10-01 14:19:05 +03:00
Avi Kivity	e99426df60	treewide: de-static namespace scope functions in headers 'static inline' is always wrong in headers - if the same header is included multiple times, and the function happens not to be inlined, then multiple copies of it will be generated. Fix by mechanically changing '^static inline' to 'inline'.	2024-10-01 14:02:50 +03:00
Avi Kivity	e9425e15b2	treewide: remove dependency on boost asio address_v4 It's not used. There's a comment mentioning it prevents some type conflict, but apparently that was fixed some time ago. Closes scylladb/scylladb#20883	2024-10-01 14:00:50 +03:00
Pavel Emelyanov	24598848a9	Merge 'virtual_tables: snapshots: include all snapshots' from Benny Halevy Use database::get_snapshot_details to get the details of all snapshots on disk, in particular those of deleted tables. Add test_snapshots_dropped_table to test listing of snapshots of a deleted table. And harden the existing test cases to use a unique snapshot tag and to delete it when the test ends. Fixes #18313 * No backport required at this time since this is rather minor UX issue that weren't hit in the field AFAIK Closes scylladb/scylladb#20869 * github.com:scylladb/scylladb: cql-pytest: test_virtual_tables: add test_snapshots_multiple_keyspaces virtual_tables: snapshots: include all snapshots	2024-10-01 13:56:13 +03:00
Gleb Natapov' via ScyllaDB development	22368b13f2	api: introduce raft stepdown REST API Also provide test.py util function to trigger it. Can be useful for testing.	2024-10-01 12:18:49 +02:00
Avi Kivity	f5628be597	Update tools/java submodule * tools/java 5b0e274f12...b2d025fd6b (1): > build.xml: update scylla-tools license	2024-10-01 12:48:45 +03:00
Pavel Emelyanov	1dfe780457	cql: Check that CREATEing tablets/vnodes is consistent with the CLI There are two bits that control whenter replication strategy for a keyspace will use tablets or not -- the configuration option and CQL parameter. This patch tunes its parsing to implement the logic shown below: if (strategy.supports_tablets) { if (cql.with_tablets) { if (cfg.enable_tablets) { return create_keyspace_with_tablets(); } else { throw "tablets are not enabled"; } } else if (cql.with_tablets = off) { return create_keyspace_without_tablets(); } else { // cql.with_tablets is not specified if (cfg.enable_tablets) { return create_keyspace_with_tablets(); } else { return create_keyspace_without_tablets(); } } } else { // strategy doesn't support tablets if (cql.with_tablets == on) { throw "invalid cql parameter"; } else if (cql.with_tablets == off) { return create_keyspace_without_tablets(); } else { // cql.with_tablets is not specified return create_keyspace_without_tablets(); } } closes: #20088 In order to enable tablets "by default" for NetworkTopologyStrategy there's explicit check near ks_prop_defs::get_initial_tablets(), that's not very nice. It needs more care to fix it, e.g. provide feature service reference to abstract_replication_strategy constructor. But since ks_prop_defs code already highjacks options specifically for that strategy type (see prepare_options() helper), it's OK for now. There's also #20768 misbehavior that's preserved in this patch, but should be fixed eventually as well. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20779	2024-10-01 10:54:29 +02:00
Botond Dénes	e780a3f168	Merge 'fix regressions of building tests with cmake' from Laszlo Ersek Fix two recent regressions of the cmake build -- found this time in the test suite. We (presumably) don't build stable releases (and their tests) with CMake, so backporting these fixes appears unnecessary, even if the regressions have been ported to stable branches. @xemul @dawmd @tchaikov @tgrabiec @scylladb/scylla-maint Closes scylladb/scylladb#20854 * github.com:scylladb/scylladb: test/boost/bptree_test: fix the CMake build test/boost/auth_test: fix the CMake build	2024-10-01 11:14:19 +03:00
Kefu Chai	d484121cc8	github: add a trigger to retrigger clang-tidy with comment before this change, clang-tidy is triggered by a pull request. but there are chances that user wants to retrigger it. for jenkins jobs, user can rebuild a job manually. but for workflow, only the developers with write permission can retrigger a workflow. this is not convenient to regular contributors. so, in this change, another trigger is added, so that user can trigger the clang-tidy workflow with "/clang-tidy" command. the syntax is inspired by IRC commands. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20841	2024-10-01 11:10:54 +03:00
Ernest Zaslavsky	5a96549c86	test: add complete_multipart_upload completion tests A primitive python http server is processing s3 client requests and issues either success or error. A multipart uploader should fail or succeed (with or without retries) depending on aforementioned server response	2024-10-01 09:06:24 +03:00
Ernest Zaslavsky	3be6052786	code: s3 client error handling Handle the `finalize_upload` possible exception to abort the upload (which also can throw) and show the right error originated from the `finalize_upload`	2024-10-01 09:06:24 +03:00
Ernest Zaslavsky	6be2433b5a	code: add response parsing and error handling to the complete_multipart_upload Instead of ignoring the response for multipart upload completion start parsing it and look for a possible errors in the response body. If the error is found throw an exception	2024-10-01 09:06:24 +03:00
Ernest Zaslavsky	826cf5cd4a	code: Introduce AWS errors parsing Add a simple utility class to parse (possible) error response from AWS S3. Stay as close as possible to aws-sdk-cpp ErrorMarshaler https://github.com/aws/aws-sdk-cpp/blob/main/src/aws-cpp-sdk-core/source/client/AWSErrorMarshaller.cpp logic Also, add a tester for this new class	2024-10-01 09:06:24 +03:00
Michał Chojnowski	c77d00fd8d	index_reader: remove a piece of misguided code involved in single-partition reads This patch removes a piece of code which, according to the comment, allows for forwarding the index reader even if it was created as a single-partition reader. For single-partition reads, the input_stream used by the reader is limited to the single index page containing the partition, since reading the index file past that point would be a waste. Because of this limit, such an index reader can't be forwarded/advanced. The dubious piece of code gets around that by unsetting the stream and ensuring it will be re-created, this time without the limit, if the index is advanced. But there is no use for this. The idea of a "single-partition reader" exist as an optimization. It's illegal to forward single-partition readers, and it doesn't make sense to attempt that. (If there's a need for forwarding, just don't create a single-partition reader). I suspect this piece of code was written due to a misunderstanding. Before the previous patch in this series, when the searched partition key was the first key in its page, the index reader would scan the preceding page first, realize it made a mistake, and advance to the next, correct page. I suspect this piece of code was written to make this work. But this is, in fact, undesirable. The fact that the index reader was working like this was a performance bug. In the single-partition case there's never an inherent reason to start with the wrong page. The index logic can be corrected to always start with the right page, and that's what the previous patch in this series does. And with that, there is no need to support advancing anymore, and the dubious piece of code can be erased. We also add an assert to emphasize that advancing a single-partition reader is illegal.	2024-09-30 23:43:36 +02:00
Michał Chojnowski	bc30509523	index_reader: in single-partition reads, don't read more than one page When looking for a partition key in the index, we scan the index from the first index page which can possibly contain the key. In a single-partition read, there is never a reason to read beyond that page. After the previous patch in this series, it's guaranteed that the first key in the next page is strictly greater than the searched key. So if the searched key is greater than the last key in the first page, then it is neither in the first nor the second page -- it must be absent from the sstable. But with the current logic, we read the second index page anyway, and the realization that the key is absent happens higher in the call chain. This patch optimizes that inefficiency by immediately returning EOF if a single-partition read doesn't find the key in the first page. Returning "end of file" even though we didn't actually go beyond the end of file is hacky, but I don't see any other non-invasive way of communicating to the caller that the partition is absent. Some caller of the index could possibly assume that returning EOF proves that the searched key is greater than all keys in the sstable. I don't think any such caller exists today, but it's a possible place for confusion. Together with the previous patch in this series, this patch guarantees that a single-partition read only accesses a single index page. This fixes a weird secondary performance bug. Due to some misunderstanding in the logic, when during a single-partition read we scan two index pages, the second index page is scanned via an input_stream created without an upper I/O limit, which means that we additionally read a full read-ahead (currently: 64 kiB) past the second index page for no reason whatsoever. After this and the previous patch, a single-partition read always reads exactly one index page, so the above problem cannot occur.	2024-09-30 23:43:36 +02:00
Michał Chojnowski	6b8b7d962c	index_reader: fix unnecessary reads of preceding index pages When setting the index to position X, we first look for the first summary entry N such that N >= X. Then we load the index page preceding N and scan it for the first partition key P such that P >= X. If there is no such key in this page, then we scan the next page (starting with N) for such key. (In this case it's always the first key). For example, assume we have: summary: A C E index: A B C D E F If we look up "B" in the index, then we first locate summary entry "C", then we scan the index for B, starting from "A". This is all fine. But when we look for "C" in the index, then we do the exactly the same -- we scan the index for "C" starting from "A". This is wasteful, because we can start scanning from "C". To avoid this inefficiency, we should be looking for N > X, not N >= X. This patch fixes that. In addition, this fixes a second, weirder performance bug. Due to some misunderstanding in the logic, when during a single-partition read we scan two index pages, the second index page is scanned via an input_stream created without an upper I/O limit, which means that we additionally read a full read-ahead (currently: 64 kiB) past the second index page for no reason whatsoever. After this patch, a single-partition read always reads exactly one index page, so the above problem cannot occur.	2024-09-30 23:43:36 +02:00
Avi Kivity	fb8743b2d6	Merge 'sstables: Fix use-after-free on page cache buffer when parsing promoted index entries across pages' from Tomasz Grabiec This fixes a use-after-free bug when parsing clustering key across pages. Also includes a fix for allocating section retry, which is potentially not safe (not in practice yet). Details of the first problem: Clustering key index lookup is based on the index file page cache. We do a binary search within the index, which involves parsing index blocks touched by the algorithm. Index file pages are 4 KB chunks which are stored in LSA. To parse the first key of the block, we reuse clustering_parser, which is also used when parsing the data file. The parser is stateful and accepts consecutive chunks as temporary_buffers. The parser is supposed to keep its state across chunks. In `93482439`, the promoted index cursor was optimized to avoid fully page copy when parsing index blocks. Instead, parser is given a temporary_buffer which is a view on the page. A bit earlier, in `b1b5bda`, the parser was changed to keep shared fragments of the buffer passed to the parser in its internal state (across pages) rather than copy the fragments into a new buffer. This is problematic when buffers come from page cache because LSA buffers may be moved around or evicted. So the temporary_buffer which is a view on the LSA buffer is valid only around the duration of a single consume() call to the parser. If the blob which is parsed (e.g. variable-length clustering key component) spans pages, the fragments stored in the parser may be invalidated before the component is fully parsed. As a result, the parsed clustering key may have incorrect component values. This never causes parsing errors because the "length" field is always parsed from the current buffer, which is valid, and component parsing will end at the right place in the next (valid) buffer. The problematic path for clustering_key parsing is the one which calls primitive_consumer::read_bytes(), which is called for example for text components. Fixed-size components are not parsed like this, they store the intermediate state by copying data. This may cause incorrect clustering keys to be parsed when doing binary search in the index, diverting the search to an incorrect block. Details of the solution: We adapt page_view to a temporary_buffer-like API. For this, a new concept is introduced called ContiguousSharedBuffer. We also change parsers so that they can be templated on the type of the buffer they work with (page_view vs temporary_buffer). This way we don't introduce indirection to existing algorithms. We use page_view instead of temporary_buffer in the promoted index parser which works with page cache buffers. page_view can be safely shared via share() and stored across allocating sections. It keeps hold to the LSA buffer even across allocating sections by the means of cached_file::page_ptr. Fixes #20766 Closes scylladb/scylladb#20837 * github.com:scylladb/scylladb: sstables: bsearch_clustered_cursor: Add trace-level logging sstables: bsearch_clustered_cursor: Move definitions out of line test, sstables: Verify parsing stability when allocating section is retried test, sstables: Verify parsing stability when buffers cross page boundary sstables: bsearch_clustered_cursor: Switch parsers to work with page_view cached_file: Adapt page_view to ContiguousSharedBuffer cached_file: Change meaning of page_view::_size to be relative to _offset rather than page start sstables, utils: Allow parsers to work with different buffer types sstables: promoted_index_block_parser: Make reset() always bring parser to initial state sstables: bsearch_clustered_cursor: Switch read_block_offset() to use the read() method sstables: bsearch_clustered_cursor: Fix parsing when allocating section is retried	2024-10-01 00:02:55 +03:00
Calle Wilund	b5d167699c	commitlog: Fix buffer_list_bytes not updated correctly Fixes #20862 With the change in `60af2f3cb2` the bookkeep for buffer memory was changed subtly, the problem here that we would shrink buffer size before we after flush use said buffer's size to decrement the buffer_list_bytes value, previously inc:ed by the full, allocated size. I.e. we would slowly grow this value instead of adjusting properly to actual used bytes. Test included. Closes scylladb/scylladb#20886	2024-09-30 18:04:00 +03:00
Raphael S. Carvalho	cf58674029	replica: Fix schema change during migration cleanup During migration cleanup, there's a small window in which the storage group was stopped but not yet removed from the list. So concurrent operations traversing the list could work with stopped groups. During a test which emitted schema changes during migrations, a failure happened when updating the compaction strategy of a table, but since the group was stopped, the compaction manager was unable to find the state for that group. In order to fix it, we'll skip stopped groups when traversing the list since they're unused at this stage of migration and going away soon. Fixes #20699. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Closes scylladb/scylladb#20798	2024-09-30 17:30:38 +03:00
David Garcia	b94fbbf30c	docs: update command Removes the update command from the setup command. This is required because versions now are not strictly pinned in the poetry.lock file since Sphinx ScyllaDB Theme 1.8. Closes scylladb/scylladb#20876	2024-09-30 17:06:07 +03:00
Andrei Chekun	cdd0c0b7fc	test.py: Do not attach logs for passed tests To reduce the amount of space needed for reports, this PR will modify logs attachment in allure, so it will attach logs only for the tests that have status other than PASSED. To simplify the solution, with the current way it's not possible to switch off these logs completely. Closes scylladb/scylladb#20786	2024-09-30 14:55:55 +02:00
Kefu Chai	1c8100d3f1	test/unit: remove unused #include following headers are no longer used by this compilation unit: - "utils/managed_ref.hh" - "test/perf/perf.hh" this was identified by clang-include-cleaner. As the code is audited, we can safely remove the #include directive. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20850	2024-09-30 14:46:39 +03:00
Kefu Chai	c3be4a36af	test.py: pass "count" to re.sub() with kwarg since Python 3.13, passing count to `re.sub()` as positional argument has been deprecated. and when runnint `test.py` with Python 3.13, we have following warning: ``` /home/kefu/dev/scylladb/./test.py:1477: DeprecationWarning: 'count' is passed as positional argument args.tests = set(re.sub(r'.* List configured unit tests\n(.*)\n', r'\1', out, 1, re.DOTALL).split("\n")) ``` see also https://github.com/python/cpython/issues/56166 in order to silence this distracting warning, let's pass `count` using kwarg. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20859	2024-09-30 13:57:02 +03:00
Kefu Chai	947d9d5a97	scylla_coredump_setup: fix typos in comment these typos were identified by the codespell workflow. and fixed a syntax error along the way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20877	2024-09-30 13:29:34 +03:00
Aleksandra Martyniuk	efc7ad8547	node_ops: fix task_manager_module::get_nodes() Currently, node ops virtual task gathers its children from all nodes contained in a sum of service::topology::normal_nodes and service::topology::transition_nodes. The maps may contain nodes that are down but weren't removed yet. So, if a user requests the status of a node ops virtual task, the task's attempt to retrieve its children list may fail with seastar::rpc::closed_error. Filter out the tasks that are down in node_ops::task_manager_module::get_nodes. Fixes: #20843. Closes scylladb/scylladb#20856	2024-09-30 12:32:23 +03:00
Pavel Emelyanov	423b5a3ba7	Merge 'directories: cleanups to silence clang-tidy false alarms' from Kefu Chai clang-tidy warns: ``` Warning: /__w/scylladb/scylladb/utils/directories.cc:132:52: warning: 'path' used after it was moved [bugprone-use-after-move] 132 \| bool can_access = co_await file_accessible(path.string(), access_flags::read \| access_flags::write \| access_flags::execute); \| ^ /__w/scylladb/scylladb/utils/directories.cc:121:28: note: move occurred here 121 \| verification_error(std::move(path), "File not owned by current euid: {}. Owner is: {}", geteuid(), sd.uid); \| ^ ``` because we pass `std::move(path)` to `verification_error()`, and "then" use this variable again in this same function. this is a false alarm, but we could make it very clear to convince this tool that it's safe to do so. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#20875 * github.com:scylladb/scylladb: directories: mark verification_error() with [[noreturn]] directories: pass const ref of path to verification_error()	2024-09-30 12:02:39 +03:00
Kamil Braun	322efb54c2	Merge 'raft_group0_client: place on a #include diet' from Avi Kivity Reduce compile time and unnecessary compilations by reducing #include load. Minor refactoring, no backport. Closes scylladb/scylladb#20864 * github.com:scylladb/scylladb: raft_group0_client: uninclude "raft_group0_registry.hh" raft_group_registry: extract raft_timeout raft_group0_client: uninclude "mutation/mutation.hh" raft_group0_client: uninclude "db/system_keyspace.hh" db: system_keyspace: extract auth_version_t into its own header	2024-09-30 10:43:44 +02:00
Kefu Chai	faec71e666	directories: mark verification_error() with [[noreturn]] this helps the compiler or static analyzers do make the right decision. for instance, clang-tidy thinks a parameter like `std::move(path)` could be reused after being moved away. with this attribute, this tool should be able to tell that this never happens. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-30 12:07:15 +08:00
Kefu Chai	0ef72475fc	directories: pass const ref of path to verification_error() before this change, we pass a `path` to `verification_error()` by moving away from the original `path`. this works fine in the sense that it is correct and does not incur potential performance issues. but clang-tidy considers it a used-after-move, because it cannot tell `verification_error()` does not return at all, and believes that `path` could be accessed again after being moved away. so it warns like: ``` Warning: /__w/scylladb/scylladb/utils/directories.cc:132:52: warning: 'path' used after it was moved [bugprone-use-after-move] 132 \| bool can_access = co_await file_accessible(path.string(), access_flags::read \| access_flags::write \| access_flags::execute); \| ^ /__w/scylladb/scylladb/utils/directories.cc:121:28: note: move occurred here 121 \| verification_error(std::move(path), "File not owned by current euid: {}. Owner is: {}", geteuid(), sd.uid); \| ^ ``` in this change, instead of passing `fs::path` to `verification_error()`, we pass a `const fs::path&` to this function. because `verification_error()` is not coroutine, neither does it not pass `path` to another continuation to be scheduled. so it's perfectly fine to pass `path` to it. this change address the false alarms from clang-tidy. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-30 12:07:15 +08:00
Nadav Har'El	64c0540d02	cql-pytest: test a few small materialized views syntax issue While documenting materialized view in a new document (Refs #16569) I encountered a few questions and this patch contains tests that clarify their answer - and can later guarantee that the answer doesn't unintentionally change in the future. The questions that these tests answer are: 1. It is not allowed to filter a view on a static column (a comment on the test explains why). 2. We already tested that it's not allowed to SELECT a static column into a view. Here we add the check that "SELECT *" is also not allowed if a static column exists in the base table. 3. We check that CREATE MATERIALIZED VIEW ... WITH COMMENT='..' works. 4. We check that CREATE MATERIALIZED VIEW ... WITH COMPACT STORAGE is forbidden. 5. We check that CREATE MATERIALIZED VIEW ... WITH garbage=.. fails with a clean InvalidRequest. All these tests pass on both Scylla and Cassandra. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20873	2024-09-29 21:34:24 +03:00
Nadav Har'El	b008dabee5	test/cql-pytest: fix support for Cassandra 3 One of the design goals of the test/cql-pytest frameworks was to be able to run these tests against Cassandra. Preferably, we should be able to run most of the tests against any popular version of Cassandra, including Cassandra 3. This is admittingly a very old version, but was still maintained until just a year ago, it's the version that Scylla is most compatible with, and we can still be curious about how it worked. Until recently cql-pytest indeed worked on Cassandra 3, but it broke on some change related to tablet detection that cause our most basic fixture - "text_keyspace" - to use the Cassandra 4 feature of "auto expand". This is trivial to fix - we should just use the this_dc fixture that we already had exactly for this purpose. Fixes #20781 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20782	2024-09-29 19:36:33 +03:00
Benny Halevy	946f21bbd3	cql-pytest: test_virtual_tables: add test_snapshots_multiple_keyspaces Test snapshots listing in system.snapshots using multiple keyspaces and multiple snpashots. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-29 14:36:18 +03:00
Benny Halevy	906de3444b	virtual_tables: snapshots: include all snapshots Use database::get_snapshot_details to get the details of all snapshots on disk, in particular those of deleted tables. Add test_snapshots_dropped_table to test listing of snapshots of a deleted table. And harden the existing test cases to use a unique snapshot tag and to delete it when the test ends. Fixes scylladb/scylladb#18313 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-29 14:16:11 +03:00
Botond Dénes	a4c41755de	Update seastar submodule * ./seastar 69f88e2f...3c9c2696 (14): > core/reactor: don't check AIO block count when they are not needed > build: do not print the default value of --c++-standard in help output > json_formatter: Add tests for formatter::write > Add APIs to get group details and to change ownership of file. > scripts/perftune.py: improve a dry-run printout > build: drop the workaround for a GCC bug > cmake: Depend on libbsd if DPDK depends on it > http: clarify the ownership in the router's doxygen comment > build: check for P2582R1 support > python: introduce a python formatting CI check > addr2line: reformat with black > scripts: add pyproject.toml > json_formatter: Make formatter::write work for std::pair > README.md: use the github homepage of Ceph for Crimson Closes scylladb/scylladb#20836	2024-09-29 13:47:40 +03:00
Avi Kivity	5a470b2bfb	Merge 'scylla_raid_setup: configure SELinux file context' from Takuya ASADA On RHEL9, systemd-coredump fails to coredump on /var/lib/scylla/coredump because the service only have write acess with systemd_coredump_var_lib_t. To make it writable, we need to add file context rule for /var/lib/scylla/coredump, and run restorecon on /var/lib/scylla. Fixes #19325 Closes scylladb/scylladb#20528 * github.com:scylladb/scylladb: scylla_raid_setup: configure SELinux file context scylla_coredump_setup: fix SELinux configuration for RHEL9	2024-09-29 12:53:00 +03:00
Avi Kivity	884297ae2e	raft_group0_client: uninclude "raft_group0_registry.hh" Reduce unnecessary recompilations.	2024-09-28 17:25:11 +03:00
Avi Kivity	67cdd0d389	raft_group_registry: extract raft_timeout It is a vocabulary term that shouldn't need the registry to be visible. Extract it to a new header.	2024-09-28 17:25:03 +03:00
Avi Kivity	93afc77307	raft_group0_client: uninclude "mutation/mutation.hh" Lighten the dependency load. Some constructors and destructors are uninlined to avoid the header depending on the mutation class.	2024-09-28 16:31:53 +03:00
Avi Kivity	5d68efe0bd	raft_group0_client: uninclude "db/system_keyspace.hh" It doesn't need it apart from a forward declaration. Files that lost necessary includes are adjusted, and some users of auth_version_t are redirected to the definition outside system_keyspace.	2024-09-28 16:31:53 +03:00
Avi Kivity	df3ee94467	db: system_keyspace: extract auth_version_t into its own header Users of auth_version_t shouldn't need to include the heavyweight system_keyspace.hh.	2024-09-28 16:31:50 +03:00
Pavel Emelyanov	c17d353718	Revert "[script/pull_github_pr.sh] Check Gating status before merging" This reverts commit `fac682df7e`. Again, this patch broke maintainer workflows, it needs even more care.	2024-09-27 19:12:18 +03:00
Benny Halevy	23d6b996b8	test/pylib: scylla_cluster: set endpoint_snitch in scylla conf When `property_file` is provided, we generate a `cassandra-rackdc.properties` file, but to actually use it, `endpoint_snitch` must be set to `GossipingPropertyFileSnitch`. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#20730	2024-09-27 16:46:54 +03:00
David Garcia	4900e4b1ac	docs: update theme 1.8.1 chore: update README Closes scylladb/scylladb#20832	2024-09-27 14:35:39 +02:00
Laszlo Ersek	153279dbfa	test/boost/bptree_test: fix the CMake build Commit `4cf4b7d4ef` ("test: Move B+tree compactiont test from unit to boost", 2024-09-24) introduced the first SEASTAR_THREAD_TEST_CASE to "test/boost/bptree_test.cc" (alongside the prior BOOST_AUTO_TEST_CASEs), but missed changing the KIND of the test from BOOST to SEASTAR. Therefore we get a linker failure: > : && /usr/bin/clang++ -O2 -Xlinker --build-id=sha1 --ld-path=ld.lld > -dynamic-linker=/.../lib64/ld-linux-x86-64.so.2 > test/boost/CMakeFiles/bptree_test.dir/Dev/bptree_test.cc.o -o > test/boost/Dev/bptree_test -L$srcdir/idl/absl::headers > -Wl,-rpath,$srcdir/idl/absl::headers test/lib/Dev/libtest-lib.a > seastar/Dev/libseastar.a /usr/lib64/libxxhash.so > /usr/lib64/libboost_unit_test_framework.so.1.83.0 utils/Dev/libutils.a > -Xlinker --push-state -Xlinker --whole-archive auth/Dev/libscylla_auth.a > -Xlinker --pop-state /usr/lib64/libcrypt.so cdc/Dev/libcdc.a > compaction/Dev/libcompaction.a mutation_writer/Dev/libmutation_writer.a > -Xlinker --push-state -Xlinker --whole-archive dht/Dev/libscylla_dht.a > -Xlinker --pop-state types/Dev/libtypes.a index/Dev/libindex.a -Xlinker > --push-state -Xlinker --whole-archive locator/Dev/libscylla_locator.a > -Xlinker --pop-state message/Dev/libmessage.a gms/Dev/libgms.a > sstables/Dev/libsstables.a readers/Dev/libreaders.a > schema/Dev/libschema.a -Xlinker --push-state -Xlinker --whole-archive > tracing/Dev/libscylla_tracing.a -Xlinker --pop-state > Dev/libscylla-main.a -Xlinker --push-state -Xlinker --whole-archive > Dev/libscylla-zstd.a -Xlinker --pop-state /usr/lib64/libzstd.so > abseil/absl/strings/Dev/libabsl_cord.a > abseil/absl/strings/Dev/libabsl_cordz_info.a > abseil/absl/strings/Dev/libabsl_cord_internal.a > abseil/absl/strings/Dev/libabsl_cordz_functions.a > abseil/absl/strings/Dev/libabsl_cordz_handle.a > abseil/absl/crc/Dev/libabsl_crc_cord_state.a > abseil/absl/crc/Dev/libabsl_crc32c.a > abseil/absl/crc/Dev/libabsl_crc_internal.a > abseil/absl/crc/Dev/libabsl_crc_cpu_detect.a > abseil/absl/strings/Dev/libabsl_str_format_internal.a /usr/lib64/libz.so > service/Dev/libservice.a node_ops/Dev/libnode_ops.a > service/Dev/libservice.a node_ops/Dev/libnode_ops.a -lsystemd > raft/Dev/libraft.a repair/Dev/librepair.a streaming/Dev/libstreaming.a > replica/Dev/libreplica.a db/Dev/libdb.a mutation/Dev/libmutation.a > data_dictionary/Dev/libdata_dictionary.a cql3/Dev/libcql3.a > transport/Dev/libtransport.a cql3/Dev/libcql3.a > transport/Dev/libtransport.a lang/Dev/liblang.a > /usr/lib64/liblua-5.4.so -lm /usr/lib64/libsnappy.so.1.1.10 > abseil/absl/container/Dev/libabsl_raw_hash_set.a > abseil/absl/hash/Dev/libabsl_hash.a abseil/absl/hash/Dev/libabsl_city.a > abseil/absl/types/Dev/libabsl_bad_variant_access.a > abseil/absl/hash/Dev/libabsl_low_level_hash.a > abseil/absl/types/Dev/libabsl_bad_optional_access.a > abseil/absl/container/Dev/libabsl_hashtablez_sampler.a > abseil/absl/profiling/Dev/libabsl_exponential_biased.a > abseil/absl/synchronization/Dev/libabsl_synchronization.a > abseil/absl/debugging/Dev/libabsl_stacktrace.a > abseil/absl/synchronization/Dev/libabsl_graphcycles_internal.a > abseil/absl/synchronization/Dev/libabsl_kernel_timeout_internal.a > abseil/absl/debugging/Dev/libabsl_symbolize.a > abseil/absl/debugging/Dev/libabsl_debugging_internal.a > abseil/absl/base/Dev/libabsl_malloc_internal.a > abseil/absl/debugging/Dev/libabsl_demangle_internal.a > abseil/absl/time/Dev/libabsl_time.a > abseil/absl/strings/Dev/libabsl_strings.a > abseil/absl/strings/Dev/libabsl_strings_internal.a > abseil/absl/strings/Dev/libabsl_string_view.a > abseil/absl/base/Dev/libabsl_throw_delegate.a > abseil/absl/numeric/Dev/libabsl_int128.a > abseil/absl/base/Dev/libabsl_base.a > abseil/absl/base/Dev/libabsl_raw_logging_internal.a > abseil/absl/base/Dev/libabsl_log_severity.a > abseil/absl/base/Dev/libabsl_spinlock_wait.a -lrt > abseil/absl/time/Dev/libabsl_civil_time.a > abseil/absl/time/Dev/libabsl_time_zone.a rust/Dev/libwasmtime_bindings.a > rust/librust_combined.a utils/Dev/libutils.a seastar/Dev/libseastar.a > /usr/lib64/libboost_program_options.so /usr/lib64/libboost_thread.so > /usr/lib64/libboost_chrono.so /usr/lib64/libboost_atomic.so > /usr/lib64/libcares.so /usr/lib64/libfmt.so.10.2.1 /usr/lib64/liblz4.so > /usr/lib64/libgnutls.so -latomic /usr/lib64/libsctp.so > /usr/lib64/libprotobuf.so /usr/lib64/libyaml-cpp.so > /usr/lib64/libhwloc.so /usr/lib64/libnuma.so /usr/lib64/libxxhash.so > /usr/lib64/libcryptopp.so /usr/lib64/libdeflate.so > /usr/lib64/libboost_regex.so.1.83.0 /usr/lib64/libicui18n.so > /usr/lib64/libicuuc.so -ldl && : > ld.lld: error: undefined symbol: main > >>> referenced by > /usr/bin/../lib/gcc/x86_64-redhat-linux/14/../../../../lib64/crt1.o:(_start) > > ld.lld: error: undefined symbol: > seastar::testing::seastar_test::seastar_test(char const, char const, > int, boost::unit_test::decorator::collector_t&) > ooo referenced by bptree_test.cc > >>> > test/boost/CMakeFiles/bptree_test.dir/Dev/bptree_test.cc.o:(_GLOBAL__sub_I_bptree_test.cc) > clang++: error: linker command failed with exit code 1 (use -v to see invocation) Fix the KIND now. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-09-27 12:21:17 +02:00
Laszlo Ersek	5fa87cb1c6	test/boost/auth_test: fix the CMake build Commit `78ab1ee8b7` ("test: Add tests for `CREATE ROLE WITH SALTED HASH`", 2024-09-20) made test/boost/auth_test dependent on cql3, but didn't encode the dependency in "CMakeLists.txt": > FAILED: > test/boost/CMakeFiles/auth_test.dir/RelWithDebInfo/auth_test.cc.o > /usr/bin/clang++ -DBOOST_ALL_DYN_LINK -DFMT_SHARED > -DSCYLLA_BUILD_MODE=release -DSEASTAR_API_LEVEL=7 > -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT > -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SSTRING > -DSEASTAR_TESTING_MAIN -DXXH_PRIVATE_API > -DCMAKE_INTDIR=\"RelWithDebInfo\" -I$srcdir -I$srcdir/build/gen > -I$srcdir/seastar/include -I$srcdir/build/seastar/gen/include > -I$srcdir/build/seastar/gen/src -isystem $srcdir/abseil -isystem > $srcdir/build/rust -ffunction-sections -fdata-sections -O3 -g -gz > -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra > -Wno-error=deprecated-declarations -Wimplicit-fallthrough > -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags > -Wno-missing-field-initializers -Wno-overloaded-virtual > -Wno-unsupported-friend -Wno-enum-constexpr-conversion > -Wno-unused-parameter -ffile-prefix-map=$srcdir/build=. -march=westmere > -Xclang -fexperimental-assignment-tracking=disabled -mllvm > -inline-threshold=2500 -fno-slp-vectorize -Werror=unused-result -MD -MT > test/boost/CMakeFiles/auth_test.dir/RelWithDebInfo/auth_test.cc.o -MF > test/boost/CMakeFiles/auth_test.dir/RelWithDebInfo/auth_test.cc.o.d -o > test/boost/CMakeFiles/auth_test.dir/RelWithDebInfo/auth_test.cc.o -c > $srcdir/test/boost/auth_test.cc > $srcdir/test/boost/auth_test.cc:22:10: fatal error: 'cql3/CqlParser.hpp' > file not found > 22 \| #include "cql3/CqlParser.hpp" > \| ^~~~~~~~~~~~~~~~~~~~ > 1 error generated. State the dependency now. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-09-27 11:38:03 +02:00
Tomasz Grabiec	b5ae7da9d2	sstables: bsearch_clustered_cursor: Add trace-level logging	2024-09-27 01:25:15 +02:00
Tomasz Grabiec	8e54ecd38e	sstables: bsearch_clustered_cursor: Move definitions out of line In order to later use the formatter for the inner class promoted_index_block, which is defined out of line after cached_promoted_index class definition.	2024-09-27 01:25:15 +02:00
Tomasz Grabiec	0279ac5faa	test, sstables: Verify parsing stability when allocating section is retried	2024-09-27 01:25:15 +02:00
Tomasz Grabiec	c09fa0cb98	test, sstables: Verify parsing stability when buffers cross page boundary	2024-09-27 01:25:15 +02:00
Tomasz Grabiec	7670ee701a	sstables: bsearch_clustered_cursor: Switch parsers to work with page_view This fixes a use-after-free bug when parsing clustering key across pages. Clustering key index lookup is based on the index file page cache. We do a binary search within the index, which involves parsing index blocks touched by the algorithm. Index file pages are 4 KB chunks which are stored in LSA. To parse the first key of the block, we reuse clustering_parser, which is also used when parsing the data file. The parser is stateful and accepts consecutive chunks as temporary_buffers. The parser is supposed to keep its state across chunks. In `b1b5bda`, the parser was changed to keep shared fragments of the buffer passed to the parser in its internal state (across pages) rather than copy the fragments into a new buffer. This is problematic when buffers come from page cache because LSA buffers may be moved around or evicted. So the temporary_buffer which is a view on the LSA buffer is valid only around the duration of a single consume() call to the parser. If the blob which is parsed (e.g. variable-length clustering key component) spans pages, the fragments stored in the parser may be invalidated before the component is fully parsed. As a result, the parsed clustering key may have incorrect component values. This never causes parsing errors because the "length" field is always parsed from the current buffer, which is valid, and component parsing will end at the right place in the next (valid) buffer. The problematic path for clustering_key parsing is the one which calls primitive_consumer::read_bytes(), which is called for example for text components. Fixed-size components are not parsed like this, they store the intermediate state by copying data. This may cause incorrect clustering keys to be parsed when doing binary search in the index, diverting the search to an incorrect block. The solution is to use page_view instead of temporary_buffer, which can be safely shared via share() and stored across allocating section. The page_view maintains its hold to the LSA buffer even across allocating sections. Fixes #20766	2024-09-27 01:25:15 +02:00
Tomasz Grabiec	c15145b71d	cached_file: Adapt page_view to ContiguousSharedBuffer	2024-09-27 01:25:15 +02:00
Tomasz Grabiec	29498a97ae	cached_file: Change meaning of page_view::_size to be relative to _offset rather than page start Will be easier to implement ContiguousSharedBuffer API as the buffer size will be equal to _size.	2024-09-27 01:25:15 +02:00
Tomasz Grabiec	c0fa49bab5	sstables, utils: Allow parsers to work with different buffer types Currently, parsers work with temporary_buffer<char>. This is unsafe when invoked by bsearch_clustered_cursor, which reuses some of the parsers, and passes temporary_buffer<char> which is a view onto LSA buffer which comes from the index file page cache. This view is stable only around consume(). If parsing requires more than one page, it will continue with a different input buffer. The old buffer will be invalid, and it's unsafe for the parser to store and access it. Unfortunetly, the temporary_buffer API allows sharing the buffer via the share() method, which shares the underlying memory area. This is not correct when the underlying is managed by LSA, because storage may move. Parser uses this sharing when parsing blobs, e.g. clustering key components. When parsing resumes in the next page, parser will try to access the stored shared buffers pointing to the previous page, which may result in use-after-free on the memory area. In prearation for fixing the problem, parametrize parsers to work with different kinds of buffers. This will allow us to instantiate them with a buffer kind which supports sharing of LSA buffers properly in a safe way. It's not purely mechanical work. Some parts of the parsing state machine still works with temporary_buffer<char>, and allocate buffers internally, when reading into linearized destination buffer. They used to store this destination in _read_bytes vector, same field which is used to store the shared buffers. Now it's not possible, since shared buffer type may be different than temporary_buffer<char>. So those paths were changed to use a new field: _read_bytes_buf.	2024-09-27 01:24:54 +02:00
Tomasz Grabiec	93bfaf4282	sstables: promoted_index_block_parser: Make reset() always bring parser to initial state When reset() is done due to allocating section retry, it can be theoretically in an arbitrary point. So we should not assume that it finished parsing and state was reset by previous parsing. We should reset all the fields.	2024-09-27 01:23:43 +02:00
Tomasz Grabiec	ac823b1050	sstables: bsearch_clustered_cursor: Switch read_block_offset() to use the read() method To unify logic which handles allocating section retry, and thus improve safety.	2024-09-27 01:22:35 +02:00
Nadav Har'El	9af43dcd06	Merge 'Move collections stress tests from unit/ to boost/' from Pavel Emelyanov Collection stress tests include testing of B- B+- and radix trees, and those tests live in unit/ suite. There are also small corner-case tests for those collections in boost/ suite. There's an attempt to get rid of unit suite in favor of boost one, and this PR moves the collections stress testing from unit suite into their boost counterparts. refs: scylladb/qa-tasks#1655 Closes scylladb/scylladb#20475 * github.com:scylladb/scylladb: test: Move other collection-testing headers from unit to boost test: Move stress-collecton header from unit to boost test: Move B+tree compactiont test from unit to boost test: Move radix tree compactiont test from unit to boost test: Move B-tree compactiont test from unit to boost test: Move radix tree stress test from unit to boost test: Move B-tree stress test from unit to boost test: Move b+tree stress test from unit to boost test: Add bool in_thread argument to stress_collection function	2024-09-26 18:11:23 +03:00
Botond Dénes	9fe64b5d70	Merge 'Remove datadir string from table::config' from Pavel Emelyanov The datadir keeps path to directory where local sstables can be. The very same information is now kept in table's storage options (#20542). This set fixes the remaining places that still use table::config::datadir and table::dir() and removes the datadir field. Closes scylladb/scylladb#20675 * github.com:scylladb/scylladb: treewide: Remove table::config::datadir distributed_loader: Print storage options, not datadir data_dictionary: Add formatter for storage_options test: Construct table_for_tests with table storage options test: Generalize pair of make_table_for_tests helpers tests: Add helper to get snapshot directory from storage options table: snapshot_exists: Get directory from storage options table: snapshot_on_all_shards: Get directory from storage options	2024-09-26 15:26:45 +03:00
Kamil Braun	9224e48d6b	Merge 'Populate raft address map from gossiper on raft configuration change' from Gleb Natapov For each new node added to the raft config populate its ID to IP mapping in raft address map from the gossiper. The mapping may have expired if a node is added to the raft configuration long after it first appears in the gossiper. Fixes scylladb/scylladb#20600 Backport to all supported versions since the bug may cause bootstrapping failure. Closes scylladb/scylladb#20601 * github.com:scylladb/scylladb: test: extend existing test to check that a joining node can map addresses of all pre-existing nodes during join group0: make sure that address map has an entry for each new node in the raft configuration	2024-09-26 12:41:25 +02:00
Tomasz Grabiec	8aca93b3ec	sstables: bsearch_clustered_cursor: Fix parsing when allocating section is retried Parser's state was not reset when allocating section was retried. This doesn't cause problems in practice, because reserves are enough to cover allocation demands of parsing clustering keys, which are at most 64K in size. But it's still potentially unsafe and needs fixing.	2024-09-26 12:34:41 +02:00
Laszlo Ersek	ed91d35171	sstables: coroutinize sstable::load() Best viewed with "git show -b -W". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#20822	2024-09-26 13:26:22 +03:00
Lakshmi Narayanan Sreethar	7beea03196	build: cmake: link cql3 library to the service library After commit `d16ea0af`, compiling the server using cmake fails with the following error : ``` FAILED: service/CMakeFiles/service.dir/Dev/qos/service_level_controller.cc.o ... /home/Scylla/scylladb/cql3/util.hh:21:10: fatal error: 'cql3/CqlParser.hpp' file not found 21 \| #include "cql3/CqlParser.hpp" \| ^~~~~~~~~~~~~~~~~~~~ 1 error generated. ``` Fix it by linking the cql3 to the service library. Closes scylladb/scylladb#20805	2024-09-26 09:17:30 +03:00
Yaron Kaikov	fac682df7e	[script/pull_github_pr.sh] Check Gating status before merging Maintainers use scripts/pull_github_pr.sh from scylladb.git when merging PRs and before pushing to the next. We want to prevent merges from piling up on top of unstable builds. This change will check Gating's current status and notify the maintainers Related to scylladb/scylla-pkg#3644 Closes scylladb/scylladb#20742	2024-09-26 08:44:06 +03:00
Nadav Har'El	7715abfc56	Merge 'Alternator store ProvisionedThroughput' from Amnon Heiman When users create a table using the Alternator API, they can decide if the billing is PROVISIONED of PAY_PER_REQUEST. If the billing is set to PROVISIONED, they need to set the ProvisionedThroughput ReadCapacityUnits (RCU) and WriteCapacityUnits (WCU). This series adds support for getting and setting the ProvisionedThroughput. The values will be stored as table extension tags. Following how TTL is stored within the Alternator, we will use ```system:rcu_attribute``` and ```system:wcu_attribute``` for the labels. The series adds a test that sets ProvisionedThroughput and validates that it gets the value back. It was tested with both Alternator and AWS. This series is part of the effort to monitor, limit, and bill Alternator operations. New code, no need to backport. Closes scylladb/scylladb#20056 * github.com:scylladb/scylladb: docs/alternator/compatibility.md: explain the consumed capacity provisioned Add test/alternator/test_provisioned_throughput.py test/alternator/util.py: Allow override BillingMode alternator/executor.cc: Store ProvisionedThroughput	2024-09-26 01:23:17 +03:00
Avi Kivity	357168114b	cql3: statement_restrictions: use the evaluator to calculate token for constrained global index query A global index has a primary key of the form (indexed_column, token, partition_key_column..., clustering_key_column...) The primary key columns are used to point at the base table row, and the token (computed as token(partition_key_column...) is used to maintain sort order. The query planner has an optimization: if the partition key is fully constrained to a unique value, then we compute the token from the partition key and use that to seek directly into the clustering row range for that base table partition. If the clustering key is also partially constrained, it is used to refine the index clustering key. Currently, this optimization is implemented as a hack: the partition key is extracted from the prepared statement + query options in get_global_index_token_clustering_ranges(), then used to calculate the token, which is then substituted in the expression passed to get_single_column_clustering_bounds() (the expression is shared across all running queries, so this is quite dangerous). We simplify the whole thing: - Let prepare_index_global() recognize that if the partition key is not fully constrained, then there is no way that we'll be able to compute the token (as it needs all partition key columns). Since the token is the first clustering key column of the index table, we can truncate it to length zero and bail out. - Otherwise, the partition key is fully constrained. We refactor the predicate (pk1 = :a AND pk2 = :b) to (pk1, pk2) := (:a, :b). We then pass expressions representing the partition key to the token function, ending up with token(:a, :b). We then substitute this expression into (*_idx_tbl_ck_prefix)[0], which computes the first clustering key column for the index table. - Remove the runtime component in get_global_index_clustering_ranges(). Note this include the early return if the partition key wasn't fully constrained (though the comment only mentions over-constraining), and the token computation, which is now done by evaluate(). Closes scylladb/scylladb#20733	2024-09-25 22:48:16 +03:00
Yaron Kaikov	d164fd45bc	install-dependencies.sh: update node_exporter to 1.8.2 Update node_exporter to 1.8.2 Fixes: #18493 Closes scylladb/scylladb#20254 [avi: regenerate frozen toolchain, with new clang in https://devpkg.scylladb.com/clang/clang-18.1.8-Fedora-40-aarch64.tar.gz https://devpkg.scylladb.com/clang/clang-18.1.8-Fedora-40-x86_64.tar.gz new clang regenerated due to new packaging format (`f6fe4d9e73`) and some other minor changes.]	2024-09-25 18:42:25 +03:00
Gleb Natapov	9e4cd32096	test: extend existing test to check that a joining node can map addresses of all pre-existing nodes during join	2024-09-25 17:10:09 +03:00
Kamil Braun	7d8f1d251a	Merge 'Mark node as being replaced earlier' from Gleb Natapov Before `17f4a151ce` the node was marked as been replaced in join_group0 state, before it actually joins the group0, so by the time it actually joins and starts transferring snapshot/log no traffic is sent to it. The commit changed this to mark the node as being replaced after the snapshot/log is already transferred so we can get the traffic to the node while it sill did not caught up with a leader and this may causes problems since the state is not complete. Mark the node as being replaced earlier, but still add the new node to the topology later as the commit above intended. Fixes: scylladb/scylladb#20629 Need to be backported since this is a regression Closes scylladb/scylladb#20743 * github.com:scylladb/scylladb: test: amend test_replace_reuse_ip test to check that there is no stale writes after snapshot transfer starts topology coordinator:: mark node as being replaced earlier topology coordinator: do metadata barrier before calling finish_accepting_node() during replace	2024-09-25 15:46:12 +02:00
Kamil Braun	09c68c0731	service: raft: fix rpc error message What it called "leader" is actually the destination of the RPC. Trivial fix, should be backported to all affected versions. Closes scylladb/scylladb#20789	2024-09-25 15:46:37 +03:00
Kefu Chai	d5b348460f	config: do not provide default value for set_value() and friends before this change, `config_file::set_value()` and `config_file::set_value_on_all_shards()` provide default value for `config_source`. but the default value is never used -- we alway specify the `source_source` when calling `set_value_on_all_shards()`. so in hope to improve the readability, the default value is removed. so, for example, one can figure out when `config_source::Internal` is used with less efforts. despite that `config_file::set_value()` is not used in the tree. for the sake of completeness, its default value is also dropped. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20728	2024-09-25 15:45:42 +03:00
Anna Stuchlik	8145109120	doc: add OS support for version 6.2 This commit adds the OS support for version 6.2. In addition, it removes support for 6.0, as the policy is only to include information for the supported versions, i.e., the two latest versions. Fixes https://github.com/scylladb/scylladb/issues/20804 Closes scylladb/scylladb#20806	2024-09-25 15:39:23 +03:00
Pavel Emelyanov	ae76481444	Merge 'treewide: add "table" parameter to "backup" API ' from Kefu Chai with this parameter, "backup" API can backup the given table, this enables it to be a drop-in replacement of existing rclone API used by scylla manager. Fixes https://github.com/scylladb/scylladb/issues/20636 --- this change is a part of the efforts to bring the native backup/restore to scylla, no need to backprt. Closes scylladb/scylladb#20661 * github.com:scylladb/scylladb: backup_task: fix the indent treewide: add "table" parameter to "backup" API	2024-09-25 10:53:38 +03:00
Takuya ASADA	f6fe4d9e73	toolchain: fix broken INSTALL_FROM mode We found that --clang-build-mode INSTALL_FROM tries to rebuild clang even we use an archive of prebuilt image. Seems like it is because ninja detected changes on standard library headers, which updated when we build new frozen toolchain container image. To avoid such unnecessary rebuild, we should stop archive whole clang build directory, we should archive install image instead. To do so, we can use "DESTDIR=<sysroot dir> ninja install-distribution-stripped", and archive sysroot dir as clang archive. Fixes #20421 Closes scylladb/scylladb#20422	2024-09-25 10:48:56 +03:00
Anna Stuchlik	da8047a834	doc: add an intro to the Features page This commit modifies the Features page in the following way: - It adds a short introduction and descriptions to each listed feature. - It hides the ToC (required to control and modify the information on the page, e.g., to add descriptions, have full control over what is displayed, etc.) - Removes the info about Enterprise features (following the request not to include Enterprise info in the OSS docs) Fixes https://github.com/scylladb/scylladb/issues/20617 Blocks https://github.com/scylladb/scylla-enterprise/pull/4711 Closes scylladb/scylladb#20635	2024-09-25 08:50:21 +03:00
Aleksandra Martyniuk	3195ebd04e	node_ops: make node_ops tasks type more human-friendly Currently, node ops tasks type is retrieved from topology_request without any change. Use respective node operation name instead. Closes scylladb/scylladb#20671	2024-09-25 08:49:34 +03:00
Kamil Braun	69b4769418	test: fix `topology_custom/test_raft_recovery_stuck` flakiness The test performs consecutive schema changes in RECOVERY mode. The second change relies on the first. However the driver might route the changes to different servers and we don't have group 0 to guarantee linearizability. We must rely on the first change coordinator to push the schema mutations to other servers before returning, but that only happens when it sees other servers as alive when doing the schema change. It wasn't guaranteed in the test. Fix this. Fixes scylladb/scylladb#20791 Should be backported to all branches containing this test to reduce flakiness. Closes scylladb/scylladb#20792	2024-09-25 08:45:37 +03:00
Kefu Chai	54858b8242	backup_task: fix the indent Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-25 09:11:26 +08:00
Kefu Chai	d663b6c13b	treewide: add "table" parameter to "backup" API with this parameter, "backup" API can backup the given table, this enables it to be a drop-in replacement of existing rclone API used by scylla manager. in this change: * api/storage_service: add "table" parameter to "backup" API. * snapshot_ctl: compose the full path of the snapshot directory in `snapshot_ctl::start_backup`. since we have all the information for composing the snapshot directory, and what the `backup_task_impl` class is interested is but the snapshot directory, we just pass the path to it instead the individual components of the directory. * backup_task_impl: instead of scan the whole keyspace recursively, only scan the specified snapshot directory. Fixes scylladb/scylladb#20636 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-25 09:11:26 +08:00
Avi Kivity	d16ea0afd6	Merge 'cql3: Extend DESC SCHEMA by auth and service levels' from Dawid Mędrek Auth has been managed via Raft since Scylla 6.0. Restoring data following the usual procedure (1) is error-prone and so a safer method must have been designed and implemented. That's what happens in this PR. We want to extend `DESC SCHEMA` by auth and service levels to provide a safe way to backup and restore those two components. To realize that, we change the meaning of `DESC SCHEMA WITH INTERNALS` and add a new "tier": `DESC SCHEMA WITH INTERNALS AND PASSWORDS`. * `DESC SCHEMA` -- no change, i.e. the statement describes the current schema items such as keyspaces, tables, views, UDTs, etc. * `DESC SCHEMA WITH INTERNALS` -- does the same as the previous tier and also describes auth and service levels. No information about passwords is returned. * `DESC SCHEMA WITH INTERNALS AND PASSWORDS` -- does the same as the previous tier and also includes information about the salted hashes corresponding to the passwords of roles. To restore existing roles, we extend the `CREATE ROLE` statement by allowing to use the option `WITH SALTED HASH = '[...]'`. --- Implementation strategy: * Add missing things/adjust existing ones that will be used later. * Implement creating a role with salted hash. * Add tests for creating a role with salted hash. * Prepare for implementing describe functionality of auth and service levels. * Implement describe functionality for elements of auth and service levels. * Extend the grammar. * Add tests for describe auth and service levels. * Add/update documentation. --- (1): https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/backup-restore/restore.html In case the link stops working, restoring a schema was realised by managing raw files on disk. Fixes scylladb/scylladb#18750 Fixes scylladb/scylladb#18751 Fixes scylladb/scylladb#20711 Closes scylladb/scylladb#20168 * github.com:scylladb/scylladb: docs: Update user documentation for backup and restore docs/dev: Add documentation for DESC SCHEMA test: Add tests for describing auth and service levels cql3/functions/user_function: Remove newline character before and after UDF body cql3: Implement DESCRIBE SCHEMA WITH INTERNALS AND PASSWORDS auth: Implement describing auth auth/authenticator: Add member functions for querying password hash service/qos/service_level_controller: Describe service levels data_dictionary: Remove keyspace_element.hh treewide: Start using new overloads of describe treewide: Fix indentation in describe functions treewide: Return create statement optionally in describe functions treewide: Add new describe overloads to implementations of data_dictionary::keyspace_element treewide: Start using schema::ks_name() instead of schema::keyspace_name() cql3: Refactor `description` cql3: Move description to dedicated files test: Add tests for `CREATE ROLE WITH SALTED HASH` cql3/statements: Restrict CREATE ROLE WITH SALTED HASH auth: Allow for creating roles with SALTED HASH types: Introduce a function `cql3_type_name_without_frozen()` cql3/util: Accept std::string_view rather than const sstring&	2024-09-24 21:44:32 +03:00
Tomasz Grabiec	bca8258150	Merge 'tablet: Fix single-sstable split when attaching new unsplit sstables' from Raphael "Raph" Carvalho To fix a race between split and repair here `c1de4859d8`, a new sstable generated during streaming can be split before being attached to the sstable set. That's to prevent an unsplit sstable from reaching the set after the tablet map is resized. So we can think this split is an extension of the sstable writer. A failure during split means the new sstable won't be added. Also, the duration of split is also adding to the time erm is held. For example, repair writer will only release its erm once the split sstable is added into the set. This single-sstable split is going through run_custom_job(), which serializes with other maintenance tasks. That was a terrible decision, since the split may have to wait for ongoing maintenance task to finish, which means holding erm for longer. Additionally, if split monitor decides to run split on the entire compaction group, it can cause single-sstable split to be aborted since the former wants to select all sstables, propagating a failure to the streaming writer. That results in new sstable being leaked and may cause problems on restart, since the underlying tablet may have moved elsewhere or multiple splits may have happened. We have some fragility today in cleaning up leaked sstables on streaming failure, but this single-sstable split made it worse since the failure can happen during normal operation, when there's e.g. no I/O error. It makes sense to kill run_custom_job() usage, since the single-sstable split is offline and an extension of sstable writing, therefore it makes no sense to serialize with maintenance tasks. It must also inherit the sched group of the process writing the new sstable. The inheritance happens today, but is fragile. Fixes #20626. Closes scylladb/scylladb#20737 * github.com:scylladb/scylladb: tablet: Fix single-sstable split when attaching new unsplit sstables replica: Fix tablet split execute after restart	2024-09-24 19:46:11 +02:00
Abhinav	36d68ec955	raft topology: add error for removal of non-normal nodes In the current scenario, We check if a node being removed is normal on the node initiating the removenode request. However, we don't have a similar check on the topology coordinator. The node being removed could be normal when we initiate the request, but it doesn't have to be normal when the topology coordinator starts handling the request. For example, the topology coordinator could have removed this node while handling another removenode request that was added to the request queue earlier. This commit intends to fix this issue by adding more checks in the enqueuing phase and return errors for duplicate requests for node removal. This PR fixes a bug. Hence we need to backport it. Fixes: scylladb/scylladb#20271 Closes scylladb/scylladb#20500	2024-09-24 16:11:19 +02:00
Botond Dénes	24ac408a08	Revert "[script/pull_github_pr.sh] Check Gating status before merging" This reverts commit `ec0bb42b45`. This patch broke maintainer workflows, it needs more work before it can land.	2024-09-24 16:53:02 +03:00
Artsiom Mishuta	c07306582b	test.py: deselect remove_data_dir_of_dead_node event Deselect remove_data_dir_of_dead_node event from test_random_failures due to issue scylladb/scylladb#20751 Closes scylladb/scylladb#20790	2024-09-24 14:49:00 +02:00
Dawid Mędrek	1ef51be1d7	docs: Update user documentation for backup and restore We update the relevant articles addressing backing-up and restoring the schema by specifying that the user performing it must be a superuser. We also update the required version of cqlsh. Additionally, we add an article covering the fundamental information on `DESCRIBE SCHEMA`.	2024-09-24 14:21:15 +02:00
Dawid Mędrek	5e1d7f109a	docs/dev: Add documentation for DESC SCHEMA We add documentation for developers addressing `DESCRIBE SCHEMA`. It covers the following aspects of it: * motivation, * synopsis of the solution, * implementation of the solution, as well as a few subsections explaining the details: * restoring process and its side effects, * restoring roles with passwords, * list of statements generated by `DESC SCHEMA` with examples, * implementation details.	2024-09-24 14:18:01 +02:00
Dawid Mędrek	d42f1604ad	test: Add tests for describing auth and service levels We add tests verifying the following features work correctly: * describing auth: roles, role grants, granting permissions on resources, * describing service levels: creating them and attaching to roles.	2024-09-24 14:18:01 +02:00
Dawid Mędrek	10d13f541b	cql3/functions/user_function: Remove newline character before and after UDF body We remove newline characters that are printed before and after a UDF's body. This way, we want to keep the create statement as close to what was actually provided as possible. Although there should be no semantic differences with or without the newline characters, it's a lot more convenient in testing when they're not present. Fixes scylladb/scylladb#20711	2024-09-24 14:18:01 +02:00
Dawid Mędrek	be851cef10	cql3: Implement DESCRIBE SCHEMA WITH INTERNALS AND PASSWORDS When executing `DESC SCHEMA WITH INTERNALS`, Scylla now also returns statements that can be used to recreate service levels and restore the state of auth. That encompasses granting roles and permissions as well as attaching service levels to roles. If the additional parameter `WITH PASSWORDS` is provided, the statements corresponding to recreating roles in the system will also contain the stored salted hashes.	2024-09-24 14:18:01 +02:00
Dawid Mędrek	2a27d4b4d6	auth: Implement describing auth We introduce a function `describe_auth()` in `auth::service` responsible for producing a sequence of descriptions whose corresponding CQL statement can be used to restore the state of auth.	2024-09-24 14:17:58 +02:00
Nadav Har'El	b70ab7bd64	test/boost: add README.md Add a README.md in test/boost, giving a short introduction to what this directory is and what kind of tests it contains, and how to run individual tests. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20550	2024-09-24 15:16:55 +03:00
Pavel Emelyanov	39dc340424	test: Move other collection-testing headers from unit to boost Simple and straightforward. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:42:13 +03:00
Pavel Emelyanov	f0d60c2b4d	test: Move stress-collecton header from unit to boost Now all its users are in boost suite. Once moved, the stress_collection() function no longer runs in seastar thread, and the in_thread argument is removed while the function is moved. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:42:13 +03:00
Pavel Emelyanov	4cf4b7d4ef	test: Move B+tree compactiont test from unit to boost This time the boost test needs to stop being pure-boost test, since bptree compaction test case needs to run in seastar thread. Other collection tests are already such, not bptree_test joins the party. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:42:13 +03:00
Pavel Emelyanov	d1f727669c	test: Move radix tree compactiont test from unit to boost No surprises here, just move the code and hard-code default args. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:42:13 +03:00
Pavel Emelyanov	bdcf965318	test: Move B-tree compactiont test from unit to boost This test must run in seastar thread, so put it in seastar-thread test case, fortunately btree test allows that. Just like its stress peer, this test also has two invocations from suite, so make it two distinct test cases as well. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:42:13 +03:00
Pavel Emelyanov	328b5b71d7	test: Move radix tree stress test from unit to boost Just move the code. Test "scale" is also taken from default unit test arguments. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:42:13 +03:00
Pavel Emelyanov	023cc99514	test: Move B-tree stress test from unit to boost This also moves the code, but takes into account the stress test had two invovations with suite options -- small and large. Inherit both with two distinct test cases. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:42:12 +03:00
Pavel Emelyanov	72cb835c1e	test: Move b+tree stress test from unit to boost Just move the code. And hard-code the "scale" (i.e. -- number of keys and iterations) from default arguments of the unit test. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:31:33 +03:00
Pavel Emelyanov	f0526bf6a4	test: Add bool in_thread argument to stress_collection function This code is going to be shared between seastar thread and boost tests, temporarily. So not to yield in pure boost test, add the switch. It will be removed really soon. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-24 13:31:33 +03:00
Tomasz Grabiec	bd6eeb4730	Merge 'Separate schema merging logic' from Marcin Maliszkiewicz This patch doesn't yet change how schema merging works but it prepares the ground for it by simplifying the code and separating merging logic into its own unit. It consists of: - minor cleanups of unused code - moving code into separate file - simplifying merge_keyspaces code More detailed explanation in per commit messages. Relates scylladb/scylladb#19153 Closes scylladb/scylladb#19687 * github.com:scylladb/scylladb: db: schema_applier: simplify merge_keyspaces function db: schema_applier: remove unnecessary read in merge_keyspaces db: schema_tables: move scylla specific code into create keyspace function db: move schema merging code into a separate unit db: schema_tables: export some schema management functions replica: remove unused table_selector forward declaration db: remove unused flush arg from do_merge_schema func db: remove unused read_arg_values function	2024-09-24 11:43:06 +02:00
Michał Jadwiszczak	d7945eea2a	docs/dev/service_levels: replace `unspecified` workload type with `NULL` `unspecified` workload type is an internal value and it's not exposed to user via CQL. Default value for workload type from user's perspective is `NULL`. Fixes scylladb/scylladb#20780	2024-09-24 11:43:29 +03:00
Yaron Kaikov	ec0bb42b45	[script/pull_github_pr.sh] Check Gating status before merging Maintainers use scripts/pull_github_pr.sh from scylladb.git when merging PRs and before pushing to the next. We want to prevent merges from piling up on top of unstable builds. This change will check Gating's current status and notify the maintainers Related to https://github.com/scylladb/scylla-pkg/issues/3644 Closes scylladb/scylladb#20742	2024-09-24 08:39:47 +03:00
Pavel Emelyanov	9fd8eba3ec	proxy: Don't keep truncate timeout as optional argument Because it is never such -- the only caller of truncate_blocking() always knows the timeout it want this method to use. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20620	2024-09-24 08:25:54 +03:00
Pavel Emelyanov	d64529f370	Merge 'sstables/sstables.hh: Remove unused forward declarations' from Nikos Dragazis Code cleanup, no backport needed. Closes scylladb/scylladb#20767 * github.com:scylladb/scylladb: sstables: Remove forward declaration for random_access_reader sstables: Remove forward declaration for metadata_collector sstables: Remove forward declaration for sstables_manager sstables: Remove forward declaration for sstable_writer_v2 sstables: Remove forward declaration for key	2024-09-24 07:44:56 +03:00
Andrei Chekun	da2397005b	test.py: Remount cgroup before changing files ownership Change order of functions: firstly remount, then change ownership for cgroup. It was not failing before because with privileged mode, it will mount cgroups as RW, but it's better to have this check if behavior will change. Closes scylladb/scylladb#20676	2024-09-24 07:27:24 +03:00
Kefu Chai	657ea95f4c	main: coroutinize read_config() for better readability. read_config() is not on the critical path, so the performance degradation caused by C++20 couroutine is neglectable. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20694	2024-09-24 06:30:34 +03:00
Gleb Natapov	1213f02a5a	test: skip test_lwt_semaphore::test_cas_semaphore in aarch64 debug mode The test configures write timeout to much smaller value to make the test run faster since for some writes sleep is inserted to hit the timeout, but it makes aarch64 debug flaky since timeout happens when it should not because of a natural slowness. Fixes scylladb/scylladb#20515 Closes scylladb/scylladb#20744	2024-09-23 20:46:55 +02:00
Avi Kivity	5c329e3db0	Merge 'Put sstables::test class on a diet' from Pavel Emelyanov This one is aimed at giving tests the ability to call private methods of class sstable. Some of the wrappers in the test class wrap public methods and can be removed. Closes scylladb/scylladb#20614 * github.com:scylladb/scylladb: test: Remove sstables::test::binary_search() test: Remove sstables::test::move_summary() test: Remove sstables::test::read_toc() test: Remove sstables::test::get_summary() test: Remove sstables::test::get_statistics() test: Remove sstables::test::data_read()	2024-09-23 21:40:40 +03:00
Yaniv Michael Kaul	26f2cbdfe2	optimized_clang.sh: compile with -march Add for both x86_64 compilation flags for clang, to get it compile with newer arch x86_64-v3 for x86 and ARM 8.2 level for aarch64. Tested to compile fine with both clang 18.1.6 and 18.1.8. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Closes scylladb/scylladb#20682	2024-09-23 17:40:20 +03:00
Paweł Zakrzewski	16dd58fb0d	cql3: respect the user-defined page size in aggregate queries This change allows the user to fully set the page size for the query. There's still an internal hard-limit of 1MB anyway, so there's no need to limit it to our default value (because using a larger page size might be a query optimization sometimes) Fixes #20612 Closes scylladb/scylladb#20692	2024-09-23 16:31:21 +03:00
Botond Dénes	64ed3f80c7	Merge 'Coroutinize sstable_directory::remove_unshared_sstables()' from Pavel Emelyanov This one is pretty simple ``` return do_with(std::move(data), [] { toss_data(data); return remove(std::move(data)); }); ``` it doesn't really need to do_with() since "toss_data" is non-preemptive. Still, convert it into ``` toss_data(data); co_await remove(std::move(data)); ``` Closes scylladb/scylladb#20479 * github.com:scylladb/scylladb: sstables: Restore indentation after previous patch sstables: Coroutinize remove_unshared_sstables()	2024-09-23 16:15:46 +03:00
Kefu Chai	1aa030a8cd	docs: explain precedence of configure options to explain for instance which setting takes effect if both command line options and `scylla.yaml` configures the same parameter. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20696	2024-09-23 16:12:44 +03:00
Yaniv Michael Kaul	85c0bb7ff4	optimized_clang.sh: add missing symbolic links to clang (for ccache) The removal of clang removes the symblic links ccache uses to mask itself as clang/clang++ Manually add them back, so ccache can work. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Fixes: https://github.com/scylladb/scylladb/issues/20490 Closes scylladb/scylladb#20491	2024-09-23 15:55:22 +03:00
Nikos Dragazis	1e4b67dd8a	sstables: Remove forward declaration for random_access_reader The sstables header contains a forward declaration for `random_access_reader`. This was introduced in `75dc7b799e` for no obvious reason. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-23 15:28:43 +03:00
Nikos Dragazis	83ccd5bcca	sstables: Remove forward declaration for metadata_collector The sstables header contains a forward declaration for `metadata_collector`. This was introduced in `2d6608bb88` for the return value of the `sstable_writer::get_metadata_collector()`. This function was later removed in `9e7144f719` but the forward declaration was left behind. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-23 15:28:43 +03:00
Nikos Dragazis	90aff33cb0	sstables: Remove forward declaration for sstables_manager This is a duplicate. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-23 15:28:43 +03:00
Nikos Dragazis	4efca437c8	sstables: Remove forward declaration for sstable_writer_v2 The sstables header contains a forward declaration for `sstable_writer_v2`. This was introduced in `fed5b73147` but never used. It is probably a leftover from a previous revision of the patchset. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-23 15:28:43 +03:00
Nikos Dragazis	fda98ba9f6	sstables: Remove forward declaration for key The sstables header contains a forward declaration for `key`. This was introduced in `198f55dc5c` for a reference parameter in `binary_search()`. The function was eventually moved to a different header in `4ed7e529db` but the forward declaration was left behind. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-23 15:28:26 +03:00
Piotr Dulikowski	d1c7e2effa	configure.py: deduplicate --out-final-name arg added in build.ninja Every time the ninja buildfile decides it needs to be updates, it calls the configure.py script with roughly the same set of flags. However, the --out-final-name flag is improperly handled and, on each reconfigure, one more --out-final-name flag is appended to the rebuild command. This is harmless because each instance of the flag will specify the same parameter, but slightly annoying because it bloats the generated file and the duplicated flags show up in ninja's output when reconfigure runs. Fix the problem by stripping the --out-final-name flags from the set of the flags passed to the configure.py before forwarding them to the reconfigure rule. Closes scylladb/scylladb#20731	2024-09-23 15:05:13 +03:00
Dawid Mędrek	90ce86930a	auth/authenticator: Add member functions for querying password hash We add new member functions to the interface of `auth::authenticator` responsible for querying the password hash corresponding to a given role. One method indicates whether a given authenticator uses password hashes, while the other queries them or throws an exception password hashes are not used. The rationale for extending the interface of authenticator is to be able to access salted hashes from other parts of auth. We will need them in an upcoming commit responsible for describing auth.	2024-09-23 13:55:52 +02:00
Dawid Mędrek	6517ca8920	service/qos/service_level_controller: Describe service levels We implement a member function responsible for producing instances of `cql3::description` that can be used to restore service levels.	2024-09-23 13:55:49 +02:00
Kefu Chai	40f2d4c988	build: cmake: drop scylla-jmx from the build in `3cd2a61736`, we dropped scylla-jmx from the build. but didn't update the CMake building system accordingly, this broke the CMake build, as the dependencies pointing to jmx cannot be found or fulfilled. in this change, we remove all references to jmx in the CMake build. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#20736	2024-09-23 14:20:42 +03:00
Marcin Maliszkiewicz	2df8eefd67	db: schema_applier: simplify merge_keyspaces function - removes uneccesary temporary sets/vectors - removes auto&& - moves return value instead of copying - instead adds diff references to keep readability - create and alter logic is almost the same, now it's visible better	2024-09-23 12:01:36 +02:00
Marcin Maliszkiewicz	7225538845	db: schema_applier: remove unnecessary read in merge_keyspaces read_schema_partition_for_keyspace() is already called for every changing keyspace by get_schema_complete_view() and stored in _after field so we can reuse this data.	2024-09-23 12:01:36 +02:00
Marcin Maliszkiewicz	f49822f78d	db: schema_tables: move scylla specific code into create keyspace function Since extract_scylla_specific_keyspace_info() was always coupled with create_keyspace_from_schema_partition() there is no value in separating them. By moving first into the latter we: - reduce number of exported functions - simplify arguments of create_keyspace_from_schema_partition - simplify caller's code	2024-09-23 12:01:36 +02:00
Marcin Maliszkiewicz	9792d720c9	db: move schema merging code into a separate unit It's mostly self containted and it's easier to maintain reasonably sized files. Also splitting better shows boundaries between schema and schema merging code.	2024-09-23 12:01:36 +02:00
Marcin Maliszkiewicz	208050f190	db: schema_tables: export some schema management functions In subseqent commits schema merging code will be separated from db/schema_tables.cc but code which manages schema will remain intact. So those two translation units will share some amount of code. It's similar case as with replica/database.cc which creates schema on startup, it calls functions from db/schema_tables.cc. Struct qualified_name got moved to header as it's used as read_table_mutations() argument.	2024-09-23 12:01:36 +02:00
Marcin Maliszkiewicz	258ffbd126	replica: remove unused table_selector forward declaration	2024-09-23 12:01:36 +02:00
Marcin Maliszkiewicz	4630864b58	db: remove unused flush arg from do_merge_schema func	2024-09-23 12:01:36 +02:00
Marcin Maliszkiewicz	4cce9c8b5a	db: remove unused read_arg_values function	2024-09-23 12:01:36 +02:00
Nadav Har'El	6496eab5ee	Merge 'Rename Alternator batch item count metrics' from Amnon Heiman This PR addresses multiple issues with alternator batch metrics: 1. Rename the metrics to scylla_alternator_batch_item_count with op=BatchGetItem/BatchWriteItem 2. The batch size calculation was wrong and didn't count all items in the batch. 3. Add a test to validate that the metrics values increase by the correct value (not just increase). This also requires an addition to the testing to validate ops of different metrics and an exact value change. Needs backporting to allow the monitoring to use the correct metrics names. Fixes #20571 Closes scylladb/scylladb#20646 * github.com:scylladb/scylladb: alternator:test_metrics test metrics for batch item count alternator:test_metrics Add validating the increased value alternator: Fix item counting in batch operations Alterntor rename batch item count metrics	2024-09-23 10:13:07 +03:00
Kefu Chai	2014d1c0cb	cql3: drop workaround for castas_fctn_simple() now that `e13a584ab7` has been merged, and our toolchain is based on the fedora 40 on 20240710, which should include this change. so let's drop the workaround from `51d09e6a` Refs #18508 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20750	2024-09-22 19:59:10 +03:00
Kefu Chai	fdc8773278	test/scylla_gdb: get table::_schema raw pointer with lw_shared_ptr This commit addresses an issue where accessing the raw pointer of the schema instance within `table::_schema` using `table.schema._p` was unreliable. before this change, `_p` was of type `lw_shared_ptr_counter_base`, a type-erased smart pointer, preventing direct casting to the underlying schema pointer. but we still cast it to `schema` anyway. this led to a gdb.MemoryError when dereferencing the deduced pointer: but the type of `_p` is `lw_shared_ptr_counter_base`, which is a type erased smart pointer, and it cannot be casted directly to the under pointer pointing to a `schema` instance. this results in: ``` Traceback (most recent call last): File "/home/avi/scylla/test/scylla_gdb/../../scylla-gdb.py", line 5554, in invoke self.print_key_type(seastar_lw_shared_ptr(schema['_clustering_key_type']).get().dereference(), 'clustering') File "/home/avi/scylla/test/scylla_gdb/../../scylla-gdb.py", line 5533, in print_key_type key_type = seastar_shared_ptr(key_type).get().dereference() ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ gdb.MemoryError: Cannot access memory at address 0x4000079656b0078 ``` when we are dereferencing the raw pointer deduced this way. in this change, we use the wrapper of `seastar_lw_shared_ptr` to safely obtain the raw pointer. * reenable this test previously disabled by `3d781c4f` tested using ```console $ SCYLLA=/home/kefu/dev/scylladb/master/build/release/scylla \ test/scylla_gdb/run -o junit_suite_name=scylla_gdb test_misc.py::test_schema ``` on an up-to-date fedora 40 installation. Refs `3d781c4f` Fixes scylladb/scylladb#20741 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20746	2024-09-22 18:30:16 +03:00
Avi Kivity	657848dcbb	cql3: statement_restrictions, expr: move restrictions-related expression utilities out of expression.cc Move all of the blatantly restriction-related expression utilities to statement_restrictions.cc. Some are so blatant as to include the word "restriction" in their name. Others are just so specialized that they cannot be used for anything else. The motivation is that further refactoring will be simplified if it can happen within the same module, as there will not be a need to prove it has no effect elsewhere. Most of the declarations are made non-public (in .cc file) to limit proliferation. A few are needed for tests or in select_statement.cc and so are kept public. Other than that, the only changes are namespace qualifications and removal of a now-duplicate definition ("inclusive"). Closes scylladb/scylladb#20732	2024-09-22 11:00:51 +03:00
Avi Kivity	3d781c4fc8	Update frozen toolchain * tools/java e505a6d3bb...5b0e274f12 (1): > Merge 'build.xml: install and use java-11 when building' from Kefu Chai Updates to clang 18.1.8 + LLVM patch to match Fedora 40. New optimized clang build generated and stored in https://devpkg.scylladb.com/clang/clang-18.1.8-x86_64.tar.gz https://devpkg.scylladb.com/clang/clang-18.1.8-aarch64.tar.gz Due to the loss of the jmx submodule, we no longer install java-11-openjdk. We add it in install-dependencies.sh here to compensate, pending a better solution. tools/java submodule updated to remove build failure where Java 8 was selected instead of Java 11. The scylla_gdb test suite was disabled due to a regression in gdb 15, which is brought in by the toolchain update [1]. [1] https://github.com/scylladb/scylladb/issues/20741.	2024-09-21 20:07:28 +03:00
Raphael S. Carvalho	38ce2c605d	tablet: Fix single-sstable split when attaching new unsplit sstables To fix a race between split and repair here `c1de4859d8`, a new sstable generated during streaming can be split before being attached to the sstable set. That's to prevent an unsplit sstable from reaching the set after the tablet map is resized. So we can think this split is an extension of the sstable writer. A failure during split means the new sstable won't be added. Also, the duration of split is also adding to the time erm is held. For example, repair writer will only release its erm once the split sstable is added into the set. This single-sstable split is going through run_custom_job(), which serializes with other maintenance tasks. That was a terrible decision, since the split may have to wait for ongoing maintenance task to finish, which means holding erm for longer. Additionally, if split monitor decides to run split on the entire compaction group, it can cause single-sstable split to be aborted since the former wants to select all sstables, propagating a failure to the streaming writer. That results in new sstable being leaked and may cause problems on restart, since the underlying tablet may have moved elsewhere or multiple splits may have happened. We have some fragility today in cleaning up leaked sstables on streaming failure, but this single-sstable split made it worse since the failure can happen during normal operation, when there's e.g. no I/O error. It makes sense to kill run_custom_job() usage, since the single-sstable split is offline and an extension of sstable writing, therefore it makes no sense to serialize with maintenance tasks. It must also inherit the sched group of the process writing the new sstable. The inheritance happens today, but is fragile. Fixes #20626. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-09-20 23:03:01 -03:00
Raphael S. Carvalho	999f1f1318	replica: Fix tablet split execute after restart let's assume there are 2 nodes, n1, n2. n1 is the coordinator. 1) n1 emits split 2) n1 and n2 complete split work 3) n1 becomes aware all replicas are ready for split 4) n2 restarts, but places split sstable into main group[1] 5) n1 executes split 6) n2 handles split completion, but see the main group is not empty [1]: During split, main group should only contain unsplit sstables. If all sstables are split, main must be empty. This is a result of replica not setting storage group to split mode on restart (using tablet map) and therefore sstables are incorrectly placed on main group. The fix is about looking at tablet map and setting group to split mode before sstables are populated into it. Refs #20626. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-09-20 22:28:09 -03:00
Avi Kivity	cd861bc788	row_cache: coroutinize do_update() do_with() makes the change a no-brainer, and besides, it's called once per huge update. Closes scylladb/scylladb#20735	2024-09-21 00:07:02 +02:00
Botond Dénes	488a372fdc	tool/scylla-nodetool: status: reorder endpoint calls to match old nodetool Old nodetool requested `/storage_service/tokens_endpoing` first, then `/storage_service/host_id`, while the native nodetool did it in reverse order. Most of the time this is inconsequential but there is an edge case when a node's IP address is changed. This reversing of the order results in unexpected behavior for tests, causing noise via flaky tests. Match the order of the old nodetool so that the native nodetool exhibits the behavior expected by tests (and users too probably). Fixes: scylladb/scylladb#18693 Closes scylladb/scylladb#20615	2024-09-20 15:07:16 +02:00
Dawid Mędrek	b357307406	data_dictionary: Remove keyspace_element.hh The interface is not used anywhere anymore, so we can remove it safely. It has been replaced by custom functions for each keyspace element and `cql3::description`.	2024-09-20 14:24:54 +02:00
Dawid Mędrek	7b4f9c806c	treewide: Start using new overloads of describe We continue removing `data_dictionary::keyspace_element`. In this commit, we start using the overloads returning `cql3::description` in places where the methods specified by `data_dictionary::keyspace_element` were used.	2024-09-20 14:24:54 +02:00
Dawid Mędrek	df94e92b06	treewide: Fix indentation in describe functions After modifying new functions for generating `cql3::description`, we fix indentation in them in this commit.	2024-09-20 14:24:54 +02:00
Dawid Mędrek	86722e4cea	treewide: Return create statement optionally in describe functions We add a new parameter in functions used to generate instances of `cql3::description` for types related to situations where we might not need a create statement. An example of such a scenario could be `DESCRIBE TYPES`.	2024-09-20 14:24:54 +02:00
Dawid Mędrek	0702e93e32	treewide: Add new describe overloads to implementations of data_dictionary::keyspace_element We're removing `data_dictionary::keyspace_element`. Before we can do that, we need to substitute the existing methods used for describing keyspace elements with their new versions returning `cql3::description`. That's what happens in this commit.	2024-09-20 14:24:53 +02:00
Dawid Mędrek	39cf106151	treewide: Start using schema::ks_name() instead of schema::keyspace_name() We're going to remove the interface `data_dictionary::keyspace_element`. As `schema::keyspace_name()` is an implementation of one of the methods specified by that interface, we replace its uses by `schema::ks_name()`. `schema::keyspace_name()` was an alias for it, so no semantic change has occured.	2024-09-20 14:24:53 +02:00
Dawid Mędrek	1844c71f9a	cql3: Refactor `description` In these changes, we describe the purpose of the type and make it reusable for other parts of the code. That includes ditching the existing constructors, leaving the formatting of its fields to the user of the interface. The removed constructors have been replaced by free functions so that existing code can still use them the way it did before.	2024-09-20 14:24:53 +02:00
Dawid Mędrek	05d6794e65	cql3: Move description to dedicated files We move the declaration of `description` to dedicated files to be able to create instances of it from other parts of the code. `describe_statement.cc` has been functioning as an intermediary between objects that can be described and the end user. It will still perform that duty, but we want to let other modules be able to generate descriptions on their own, without having to share an additional layer of abstraction in form of types inheriting from `data_dictionary::keyspace_element`. Those types may not perform any other function than that and thus may be redundant. Adjusting `description` to its new purpose will happen in an upcoming commit.	2024-09-20 14:24:53 +02:00
Dawid Mędrek	78ab1ee8b7	test: Add tests for `CREATE ROLE WITH SALTED HASH`	2024-09-20 14:24:53 +02:00
Dawid Mędrek	47a5469280	cql3/statements: Restrict CREATE ROLE WITH SALTED HASH We start requiring that the user issuing `CREATE ROLE WITH SALTED HASH` be a superuser. The rationale for that is the statement directly modifies a system tables, circumventing the hashing algorithm. Additionally, we correct a possible existing problem. `_options.is_superuser` in `create_role_statement` may be an empty optional, so dereferencing it without a prior check could lead to undefined behavior in the future.	2024-09-20 14:24:53 +02:00
Dawid Mędrek	206fdf2848	auth: Allow for creating roles with SALTED HASH We introduce a way to create a role with explictly provided salted hash. The algorithm for creating a role with a password works like this: 1. The user issues a statement `CREATE ROLE <role> WITH PASSWORD = '<password>' <...>`. 2. Scylla produces a hash based on the value of `<password>`. 3. Scylla puts the produced hash in `system.roles`, in the column `salted_hash`. The newly introduced way to create a role is based on a new form of the create statement: `CREATE ROLE <role> WITH SALTED HASH = '<salted_hash>` The difference in the algorithm used for processing this statement is that we insert `<salted_hash>` into `system.roles` directly, without hashing it. The rationale for introducing this new statement is that we want to be able to restore roles. The original password isn't stored anywhere in the database (as intended), so we need to rely on the column `salted_hash`.	2024-09-20 14:24:53 +02:00
Dawid Mędrek	35a92d189e	types: Introduce a function `cql3_type_name_without_frozen()` The introduced function returns the actual name of the type represented by `abstract_type`. It circumvents name processing like wrapping a type within `frozen<>` or using Cassandra's syntax. We add the function to be able to describe UDFs in the upcoming commits that require that their arguments not be `frozen<>`. We also test the implementation.	2024-09-20 14:24:53 +02:00
Dawid Mędrek	202d866892	cql3/util: Accept std::string_view rather than const sstring&	2024-09-20 14:24:53 +02:00
Avi Kivity	61d19e4464	Update tools/java submodule * tools/java 0b4accdd5e...e505a6d3bb (1): > [C-S] Make it use DCAwareRoundRobinPolicy unless rack is provided	2024-09-20 14:49:21 +03:00
Pavel Emelyanov	b45891acd7	sstables: storage: Don't keep base directory in base class This reverts commit `44bd183187` and moves the base directory back on filesystem_storage. The mentioned commit says > so we can use the base (table) directory for > e.g. pending_delete logs, in the next patch. but "next patch" doesn't use it outside of the filesystem-storage anyway. This field doesn't make sense for S3 backend. Its "location" is not location, but a key in the system.sstables, which should rather be schema ID, not /var/lib/.../keyspace/table-uuid string. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20642	2024-09-20 11:51:04 +03:00
Andrei Chekun	bd9a73c39b	Add .idea folder to .gitignore .idea directory used by JetBrains IDE's to store data about project config Closes scylladb/scylladb#20718	2024-09-20 11:49:41 +03:00
Tomasz Grabiec	8e047e8fff	gdb: Add std::set wrapper Allows accessing std::set fields from gdb, e.g.: (gdb) python for e in std_set(_promoted_index._blocks): print(e) Closes scylladb/scylladb#20650	2024-09-20 08:24:15 +03:00
Anna Stuchlik	5da7894f70	doc: move the install-jmx instructions to a common folder This commit moves the install-jmx.rst file from the install-scylla folder to the installation-common folder. All the references to the moved document are updated. This is a follow-up to https://github.com/scylladb/scylladb/pull/17969/ Closes scylladb/scylladb#20712	2024-09-20 00:36:32 +03:00
Nadav Har'El	3499c407f7	test: avoid silly "no_mode.1" labels when running tests outside test.py For the benefit of running test.py inside CI, we recently added to test/cql-pytest and test/alternator the knowledge of which "Scylla mode" (--mode) and "run number" is running (--run_id), although these concepts are alien to these two test frameworks (remember that those test frameworks can also run tests against unknown versions of Scylla or even our competitors' implementations). One unfortunate result of this change is that now if you run a test by using pytest directly (or test/*/run) instead of test.py, for example: $ cd test/alternator $ pytest --aws test_item.py::test_basic_string_put_and_get The test's success or failure reports the ugly name test_item.py::test_basic_string_put_and_get.no_mode.1 This unnecessary "no_mode.1" come from the the default values for --mode and --run_id, respectively. But there is no reason for these silly defaults. In this patch we change these defaults to None, and when they are None, they aren't tacked onto the test's name. This patch shouldn't affect running tests through test.py, because test.py always sets the --mode and --run_id options, and doesn't leave them as the default. Fixes #20512 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20513	2024-09-20 00:36:32 +03:00
Avi Kivity	b015c85d31	Merge 'gms: inet_address: drop unused raw_addr method and modernize comperators' from Benny Halevy Drop the unused `gms::inet_address::raw_addr` method and modernize operator== and operator< as class methods * Cleanup only, no backport needed Closes scylladb/scylladb#20681 * github.com:scylladb/scylladb: gms: inet_address: modernize comparison operators gms: inet_address: drop unused raw_addr method	2024-09-20 00:36:32 +03:00
Piotr Dulikowski	7e7701d436	Merge 'cql3/statements/select_statement: `SELECT ... USING SERVICE LEVEL`' from Michał Jadwiszczak Allow to specify service level used in select statement `SELECT ... USING SERVICE LEVEL sl_name`. In OSS, this only affects statement's timeout. In case both service level and timeout are specified `SELECT ... USING SERVICE LEVEL sl_name AND TIMEOUT 1h`, the timeout has higher priority as statement's timeout. Fixes scylladb/scylladb#18471 Closes scylladb/scylladb#20523 * github.com:scylladb/scylladb: test/cql-pytest: add test for `SELECT ... USING SERVICE LEVEL` cql3/Cql.g: extend grammar to allow `SELECT ... USING SERVICE LEVEL` cql3/statements/select_statement: use service level timeout cql3/attributes: add service level name field qos/service_level_controller: add method to check if service level exists in cache	2024-09-19 18:19:23 +02:00
Pavel Emelyanov	bd720dd2da	Merge 'cql3: statement_restrictions: adapt to functional style' from Avi Kivity The statement_restrictions class started life in the object-oriented style - an object that interacts with its environment via mutators and is observed via observers. This is however not suitable for its objective: to analyze the WHERE clause, select a query plan, and partition the WHERE clause atoms to the various parts demanded by the query plan (read_command and filters). Furthermore, the object oriented style makes it hard to work with as you can only call some observers after the related mutators were called. Fix this by transforming the code info a more functional style: we call a function that returns an immutable statement_restrictions object that can only be observed. This makes it easier to further change in the future, as changes will not have to consider interaction with the environment. No backport as this is a refactoring Closes scylladb/scylladb#20672 * github.com:scylladb/scylladb: cql3: statement_restrictions: use functional style cql3: statement_restrictions: calculate the index only once cql3: statement_restrictions: make it a const object	2024-09-19 18:18:28 +03:00
Kefu Chai	8cc9d783a0	sstables/sstable_directory: document components_lister::process() for better maintainability. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20693	2024-09-19 18:11:31 +03:00
Kefu Chai	7985aa97b1	main, test: use seastar::handle_signal() instead use `seastar::handle_signal()` instead of `reactor::handle_signal()`. in a recent change in seastar (c3e826ad1197f2610138f3bcfaeb0b458f8fb799), the later was marked as deprecated in favor of the former, so let's use the recommended API. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20695	2024-09-19 18:10:07 +03:00
Kefu Chai	1fd1698a90	test: btree: use BOOST_DATA_TEST_CASE() when appropriate instead grouping tests with different parameters, let's parameterize them using `BOOST_DATA_TEST_CASE()`, simpler this way. and the tests can be more structured. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20697	2024-09-19 18:09:05 +03:00
Avi Kivity	6f7c2ce0aa	Merge 'cql_server::connection: Process rebounce message in case of multiple shard migrations' from Sergey Zolotukhin During a query execution, the query can be re-bounced to another shard if the requested data is located there. Previous implementation assumed that the shard cannot be changed after first re-bounce, however with the introduction of Tablets, data could be migrated to another shard after the query was already re-bounced, causing a failure of the query execution. To avoid this issue, the query is re-bounced as needed until it is executed on the correct shard. Fixes #15465 Closes scylladb/scylladb#20493 * github.com:scylladb/scylladb: cql_server: Add a test for multiple query msg rebounces. cql_server::connection: process: rebounce msg if needed cql_server::connection: process: co-routinize connection::process_on_shard cql_server: connection: process: fixup indentation cql_server: connection: process_on_shard: drop permit parameter transport: server: pass bounce_to_shard as foreign shared ptr cql_server: connection: process: add template concept for process_fn cql_server: move process_fn_return_type to class definition	2024-09-19 17:27:55 +03:00
Gleb Natapov	1b4c255ffd	test: amend test_replace_reuse_ip test to check that there is no stale writes after snapshot transfer starts	2024-09-19 15:24:59 +03:00
Gleb Natapov	c0939d86f9	topology coordinator:: mark node as being replaced earlier Before `17f4a151ce` the node was marked as been replaced in join_group0 state, before it actually joins the group0, so by the time it actually joins and starts transferring snapshot/log no traffic is sent to it. The commit changed this to mark the node as being replaced after the snapshot/log is already transferred so we can get the traffic to the node while it sill did not caught up with a leader and this may causes problems since the state is not complete. Mark the node as being replaced earlier, but still add the new node to the topology later as the commit above intended.	2024-09-19 15:23:48 +03:00
Gleb Natapov	644e7a2012	topology coordinator: do metadata barrier before calling finish_accepting_node() during replace During replace with the same IP a node may get queries that were intended for the node it was replacing since the new node declares itself UP before it advertises that it is a replacement. But after the node starts replacing procedure the old node is marked as "being replaced" and queries no longer sent there. It is important to do so before the new node start to get raft snapshot since the snapshot application is not atomic and queries that run parallel with it may see partial state and fail in weird ways. Queries that are sent before that will fail because schema is empty, so they will not find any tables in the first place. The is pre-existing and not addressed by this patch.	2024-09-19 15:00:27 +03:00
Benny Halevy	574a08ed96	storage_service: rebuild: warn about tablets-enabled keyspaces Until we automatically support rebuild for tablets-enabled keyspaces, warn the user about them. The reason this is not an error, is that after increasing RF in a new datacenter, the current procedure is to run `nodetool rebuild` on all nodes in that dc to rebuild the new vnode replicas. This is not required for tablets, since the additional replicas are rebuilt automatically as part of ALTER KS. However, `nodetool rebuild` is also run after local data loss (e.g. due to corruption and removal of sstables). In this case, rebuild is not supported for tablets-enabled keyspaces, as tablet replicas that had lost data may have already been migrated to other nodes, and rebuilding the requested node will not know about it. It is advised to repair all nodes in the datacenter instead. Refs scylladb/scylladb#17575 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#20375	2024-09-19 14:25:46 +03:00
Pavel Emelyanov	8487f2fd93	treewide: Remove table::config::datadir It's write-only now, all the places than wanted to know where table's storage is, already use storage_options. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:06:39 +03:00
Pavel Emelyanov	350f64c38b	distributed_loader: Print storage options, not datadir When populating keyspace on boot the dist. loader prints a debugging message with ks:cf names, state and the directory from where it picks sstables. The last one is not extremely correct, as loading sstables from S3 happens from a bucket, not directory. So it's better to print the storage options, not the datadir string. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:06:39 +03:00
Pavel Emelyanov	b2fcfdcaa9	data_dictionary: Add formatter for storage_options Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:06:39 +03:00
Pavel Emelyanov	5046cfab4b	test: Construct table_for_tests with table storage options The only place that constructs table_for_tests is make_table_for_tests helper. It can and should prepare the correct storage options, because that's the last place where the target directory is still known. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:06:39 +03:00
Pavel Emelyanov	eaad4f348b	test: Generalize pair of make_table_for_tests helpers They only differ in a way they get target directory from -- one via argument, andother from test_env. Respectively, the latter can call the former. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:06:39 +03:00
Pavel Emelyanov	d9ef9bdd3b	tests: Add helper to get snapshot directory from storage options There's a bunch of tests that check the contents of snapshot directory after creating one. Add a helper for those that gets this directory via storage options, not table config. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:06:39 +03:00
Pavel Emelyanov	a734fd5c9c	table: snapshot_exists: Get directory from storage options Similarly to snapshot_on_all_shards, the way snapshot directory is evaluated is changed to rely on storage options. Two ... assumptions are that when asking for non-local snapshot existance or for a snapshot of a virtual table, it's correct to return false instead of throwing. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:06:09 +03:00
Pavel Emelyanov	24589cf00c	table: snapshot_on_all_shards: Get directory from storage options There are several things that are changed here - The target directory for snapshot is evaluated using table directory taken from its storage options, not from config - If the storage options are not "local", the snapshot_on_all_shards is failed early, it's impossible to snapshot sstables anyway - If the storage is not configured for the obtained local options, snapshotting is skilled, because it's a virtual table that's probably not supposed to have snapshots - The late failure to snapshot non-local sstables is converted into internal error, as this functionality cannot be executed as per previous change - The target path is created using fs::path operator/ overload, not by concatenating strings (it's minor change) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-19 13:05:16 +03:00
Anna Stuchlik	cdc69b4e06	doc: enable publishing docs for branch-6.2 This commit enables publishing documentation from branch-6.2. The docs will be published as UNSTABLE (the warning about version 6.1 being unstable will be displayed). Fixes https://github.com/scylladb/scylladb/issues/20643 No backport is required. Closes scylladb/scylladb#20647	2024-09-19 09:39:58 +03:00
Anna Stuchlik	400a14eefa	doc: update the unified installer instructions This commit updates the unified installer instructions to avoid specifying a given version. At the moment, we're technically unable to use variables in URLs, so we need to update the page each release. Fixes https://github.com/scylladb/scylladb/issues/20677 Closes scylladb/scylladb#20680	2024-09-19 09:28:44 +03:00
Anna Stuchlik	aa0c95c95c	doc: fix a broken link This commit fixes a link to the Manager by adding a missing underscore to the external link. Closes scylladb/scylladb#20656	2024-09-19 09:20:20 +03:00
Calle Wilund	60f8a9f39d	database: Also forced new schema commitlog segment on user initiated memtable flush Refs #20686 Refs #15607 In #15060 we added forced new commitlog segment on user initated flush, mainly so that tests can verify tombstone gc and other compaction related things, without having to wait for "organic" segment deletion. Schema commitlog was not included, mainly because we did not have tests featuring compaction checks of schema related tables, but also because it was assumed to be lower general througput. There is however no real reason to not include it, and it will make some testing much quicker and more predictable. Closes scylladb/scylladb#20691	2024-09-19 09:00:33 +03:00
Benny Halevy	5ccdf1cf1c	gms: inet_address: modernize comparison operators Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-18 17:07:51 +03:00
Benny Halevy	38540d89a1	gms: inet_address: drop unused raw_addr method Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-18 14:21:18 +03:00
Kefu Chai	b0696bd842	test: btree: use BOOST_DATA_TEST_CASE to structure parameterized tests for better readability. and for more structured tests. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20516	2024-09-18 14:16:28 +03:00
Pavel Emelyanov	eb22c2a8c8	Merge 'reader_concurrency_semaphore: improve the diagnostics dump' from Botond Dénes * Also dump diagnostics when a read times out while active (not queued). * Add the "Trigger permit" line, containing the details of the permit which caused the diagnostics dump (by e.g. timing out). * Add the "Identified bottleneck(s)" line, containing the identified bottlenecks which lead to permits being queued. This line is missing if no such bottleneck can be identified. * Document the new features, as well as the stat dump, which was added some time ago. Example of the new dump format: ``` INFO 2024-09-12 08:09:48,046 [shard 0:main] reader_concurrency_semaphore - Semaphore reader_concurrency_semaphore_dump_reader_diganostics with 8/10 count and 106192275/32768 memory resources: timed out, dumping permit diagnostics: Trigger permit: count=0, memory=0, table=ks.tbl0, operation=mutation-query, state=waiting_for_admission Identified bottleneck(s): memory permits count memory table/operation/state 3 2 26M ./push-view-updates-2/active 3 2 16M ks.tbl1/push-view-updates-1/active 1 1 15M ks.tbl2/push-view-updates-1/active 1 0 13M ks.tbl1/multishard-mutation-query/active 1 0 12M ks.tbl0/push-view-updates-1/active 1 1 10M ks.tbl3/push-view-updates-2/active 1 1 6060K ks.tbl3/multishard-mutation-query/active 2 1 1930K ks.tbl0/push-view-updates-2/active 1 0 1216K ks.tbl0/multishard-mutation-query/active 6 0 0B ks.tbl1/shard-reader/waiting_for_admission 3 0 0B ./data-query/waiting_for_admission 9 0 0B ks.tbl0/mutation-query/waiting_for_admission 2 0 0B ks.tbl2/shard-reader/waiting_for_admission 4 0 0B ks.tbl0/shard-reader/waiting_for_admission 9 0 0B ks.tbl0/data-query/waiting_for_admission 7 0 0B ks.tbl3/mutation-query/waiting_for_admission 5 0 0B ks.tbl1/mutation-query/waiting_for_admission 2 0 0B ks.tbl2/mutation-query/waiting_for_admission 8 0 0B ks.tbl1/data-query/waiting_for_admission 1 0 0B ./mutation-query/waiting_for_admission 26 0 0B permits omitted for brevity 96 8 101M total Stats: permit_based_evictions: 0 time_based_evictions: 0 inactive_reads: 0 total_successful_reads: 0 total_failed_reads: 0 total_reads_shed_due_to_overload: 0 total_reads_killed_due_to_kill_limit: 0 reads_admitted: 1 reads_enqueued_for_admission: 82 reads_enqueued_for_memory: 0 reads_admitted_immediately: 1 reads_queued_because_ready_list: 0 reads_queued_because_need_cpu_permits: 82 reads_queued_because_memory_resources: 0 reads_queued_because_count_resources: 0 reads_queued_with_eviction: 0 total_permits: 97 current_permits: 96 need_cpu_permits: 0 awaits_permits: 0 disk_reads: 0 sstables_read: 0 ``` Fixes: https://github.com/scylladb/scylladb/issues/19535 Improvement, no backport needed. Closes scylladb/scylladb#20545 * github.com:scylladb/scylladb: docs/dev/reader-concurrency-semaphore.md: update the documentation on diagnostics dumps test/boost/reader_concurrency_semaphore_test: test the new diagnostics functionality reader_concurrency_semaphore: add bottleneck self-diagnosis to diagnosis dump reader_concurrency_semaphore: include trigger permit in diagnostic dump reader_concurrency_semaphore: propagate permit to do_dump_reader_permit_diagnostics() reader_concurrency_semaphore: use consistent exception type for timeout reader_concurrency_semaphore: dump diagnostics when non-waiting reader times out	2024-09-18 14:06:05 +03:00
Botond Dénes	1efda557b1	replica/table: query_mutations(): enter the table's async gate So the table is not dropped while the query is ongoing. query() already does this but using old-fashioned enter()+leave(), convert it to use the new RAII helper. Closes scylladb/scylladb#20583	2024-09-18 14:03:22 +03:00
Pavel Emelyanov	2f4f0eb060	Merge 'Alternator: a few RBAC fixes' from Nadav Har'El The main goal of this PR is to fix a bug (#20619) in the alternator_enforce_authorization=false setting - which didn't do its job (i.e, _don't_ check permissions) when authorization is configured in CQL but not wanted in Alternator. The series also a few smaller bugs in the code that were discovered while debugging the main issue: 1. A potential use-after-free (that didn't seem to hit us in practice) is fixed. 2. A confusing error message (that was also reported in #20619) is improved. 3. Make the alternator_enforce_authorization live-updatable. There was no reason why it shouldn't be, and as this series needs to make this flag available to more code, let's just do it properly and assume the flag is live-updatable. Because the RBAC feature has not been backported to any open-source branches, neither should these fixes. But if some private branch received a backport of the RBAC feature, it should get these fixes too. Fixes #20619. Closes scylladb/scylladb#20640 * github.com:scylladb/scylladb: alternator: make alternator_enforce_authorization live-updateable alternator: fix alternator_enforce_authorization=false alternator: improve error message when unauthenticated alternator: avoid use-after-free in RBAC	2024-09-18 14:02:09 +03:00
Kefu Chai	cb1670b79b	Update seastar submodule * seastar ec5da7a6...69f88e2f (38): > build: s/Sanitizers_COMPILER_OPTIONS/Sanitizers_COMPILE_OPTIONS > test: Update httpd test with request/reply body writing sugar > http: Add sugar to request and response body writers > utils: Add util::write_to_stream() helper > seastar-addr2line: adjust llvm termination regex > README.md: add Crimson project > rpc: conditionally use fmt::runtime() based on SEASTAR_LOGGER_COMPILE_TIME_FMT > build: check the combination of Sanitizers > tls: clear session ticket before releasing > print: remove dead code > doc/lambda-coroutine-fiasco: reword for better readability > rpc: fix compilation error caused by fmt::runtime() > tutorial: explain the use case of rethrow_exception and coroutine::exception > reactor: print more informative error when io_submit fails > README.md: note GitHub discussions > prometheus: `fmt::print` to stringstream directly > doc: add document for testing with seastar > seastar/testing: only include used headers > test: Add abortable http client test cases > http/client: Add abortable make_request() API method > http/client: Abort established connections > http/client: Handle abort source in pool wait > http/client: Add abort source to factory::make() method > http/client: Pass abort_source here and there > http/client: Idnentation fix after previous patch > http/client: Merge some continuations explicitly > signal: add seastar signal api > httpd: remove unused prometheus structs > print: use fmtlib's fmt::format_string in format() > rpc: do not use seastar::format() in rpc logger > treewide: s/format/seastar::format/ > prometheus: sanitize label value for text protocol > tests: unit test prometheus wire format > io-tester: Introduce batches to rate-based submission > io-tester: Generalize issueing request and collecting its result > io-tester: Cancel intent once > io-tester: Dont carry rps/parallelism variables over lambdas > io-tester: Simplify in-flight management The breaking changes in the seastar submodule necessitate corresponding modifications in our code. These changes must be implemented together in a single commit to maintain consistency. So that each commit is buildable. following changes are included in addition to seastar submodule update: * instead of passing a `const char` for the format string, pass a templated `fmt::format_string<...>`, this depends on the `seastar::format()` change in seastar. explicitly call `fmt::runtime()` if the format string is not a consteval expression. this depends on the `seastar::format()` change in seastar. as `seastar::format()` does not accept a plain `const char` which is not constexpr anymore. pass abort_source to `dns_connection_factory::make()`. this depends on the change in seastar, which added a `abort_source` argument to the pure virtual member function of `connection_factory::make()`. call call {fmt,seastar}::format() explicitly. this is a follow up of `3e84d43f`, which takes care of all places where we should call `fmt::format()` and `seastar::format()` explicitly to disambiguate the `format()` call. but more `format()` call made their way into the source tree after `3e84d43f`. so we need fix them as well. * include used header in tests Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Update seastar submodule Please enter the commit message for your changes. Lines starting Closes scylladb/scylladb#20649	2024-09-18 13:59:22 +03:00
Gleb Natapov	bddaf498df	group0: make sure that address map has an entry for each new node in the raft configuration ID->IP mapping is added to the raft address map when the mapping first appears in the gossiper, but it is added as expiring entry. It becomes non expiring when a node is added to raft configuration. But when a node joins those two events may be distant in time (since the node's request may sit in the topology coordinator queue for a while) and mappings may expire already from the map. This patch makes sure to transfer the mapping from the gossiper for a node that is added to the raft configuration instead of assuming that the mapping is already there.	2024-09-18 13:42:38 +03:00
Amnon Heiman	8dec292698	alternator:test_metrics test metrics for batch item count This patch adds tests for the batch operations item count. The tests validate that the metrics tracking the number of items processed in a batch increase by the correct amount. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-09-18 11:31:06 +03:00
Amnon Heiman	4d57a43815	alternator:test_metrics Add validating the increased value The `check_increases_operation` now allows override the checked metric. Additionally, a custom validation value can now be passed, which make it possible to validate the amount by which a value has changed, rather than just validating that the value increased. The default behavior of validating that values have increased remains unchanged, ensuring backward compatibility. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-09-18 11:31:06 +03:00
Amnon Heiman	905408f764	alternator: Fix item counting in batch operations This patch fixes the logic for counting items in batch operations. Previously, the item count in requests was inaccurate, it count the number of tabels in get_item and the request_items in write_items. The new logic correctly counts each individual item in `BatchGetItem` and `BatchWriteItem` requests. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-09-18 11:30:59 +03:00
Amnon Heiman	515857a4a9	Alterntor rename batch item count metrics This patch renames metrics tracking the total number of items in a batch to `scylla_alternator_batch_item_count`. It uses the existing `op` label to differentiate between `BatchGetItem` and `BatchWriteItem` operations. Ensures better clarity and distinction for batch operations in monitoring. This an example of how it looks like: # HELP scylla_alternator_batch_item_count The total number of items processed across all batches # TYPE scylla_alternator_batch_item_count counter scylla_alternator_batch_item_count{op="BatchGetItem",shard="0"} 4 scylla_alternator_batch_item_count{op="BatchWriteItem",shard="0"} 4	2024-09-18 11:20:07 +03:00
Anna Mikhlin	0c7ca284ad	mergify: add support for branch-6.2 branch-6.2 is already available, adding support for it in mergify to allow backport to this new branch. in addition, since branch 5.4 reached EOL - removing it Closes scylladb/scylladb#20669	2024-09-18 08:30:41 +03:00
Ernest Zaslavsky	924325fd25	treewide: add "prefix" parameter to backup API Allow the caller to pass the prefix when performing backup and restore Fixes scylladb/scylladb#20335 Closes scylladb/scylladb#20413	2024-09-18 08:25:00 +03:00
Calle Wilund	b789361091	commitlog: Fix assertion in oversized_alloc Fixes #20633 Cannot assert on actual request_controller when releasing permit, as the release, if we have waiters in queue, will subtract some units to hand to them. Instead assert on permit size + waiter status (and if zero, also controller value) * v2 - use SCYLLA_ASSERT Closes scylladb/scylladb#20654	2024-09-18 08:22:28 +03:00
Avi Kivity	57ab5ce313	repair: row_level: simplify repair_put_row_diff_with_rpc_stream_process_op() repair_put_row_diff_with_rpc_stream_process_op() always returns stop_iteration::no (or throws). Moreover, the return value is ignored by its only caller. Simplify by returning a plain future<>. Closes scylladb/scylladb#20610	2024-09-18 08:17:09 +03:00
Botond Dénes	d72fcb11f5	Merge 'Add new GDB commands to dump sstable index file from memory and print promoted index ' from Tomasz Grabiec Closes scylladb/scylladb#20648 * github.com:scylladb/scylladb: gdb: Introduce "scylla sstable-dump-cached-index" command gdb: Introduce "scylla sstable-promoted-index" command gdb: Fix range printer for singular ranges	2024-09-18 08:13:04 +03:00
Nadav Har'El	24fb92c8ba	Merge 'cql3: simplify runtime component of selection filtering' from Avi Kivity Most of the analysis of the WHERE clause is done in statement_restrictions. It determines what parts to use for the primary or secondary index, and what parts to use for filtering. The difficult part is that it has a very wide interface. After construction, the user must pick the correct bits from many public functions. There are subtle interactions between them that are hard to untangle. This series simplifies the interface as it is used for selection filtering. In the end, only two public functions are used, both returning expressions: one for the partition-level filtering, one for the clustering row level filtering. In the end, the WHERE clause is factored into three parts: - one part goes into the read_command of the primary or secondary index - another part (that references only partition key columns and static key columns) is used to filter entire partitions - another part (that currently references only clustering key columns and regular columns, but one day may reference other columns) is used to filter clustering rows Refactoring, no backport. Closes scylladb/scylladb#20487 * github.com:scylladb/scylladb: cql3: statement_restrictions: drop accessors for single-column key restrictions cql3: selection: adjust indentation cql3: selection: delete empty loop cql3: statement_restrictions, selection: fold multi-column restrictions into row-level filter cql3: statement_restrictions, selection: merge clustering key filter and regular columns filter cql3: statement_restrictions, selection: merge partition key filter and static columns filter cql3: selection: filter regular and static rows as a single expression each cql3: statement_restrictions: collect regular column and static column filters into single expressions cql3: selection: filter clustering key as a single expression cql3: statement_restrictions: expose filter for clustering key cql3: selection: filter partition key as a single expression cql3: statement_restrictions: expose filter for partition key cql3: statement_restrictions: remove relations used for indexing from filtering cql3: statement_restrictions: bail out of find_idx if !_uses_secondary_index cql3: statement_restrictions, modification_statement: pass correct value of check_indexes cql3: statement_restrictions: correct mismatched clustering/partition restrictions references cql3: statement_restrictions: precalculate get_column_defs_for_filtering() cql3: selection: do_filter(): push static/regular row glue to higher level	2024-09-17 22:58:24 +03:00
Piotr Dulikowski	cc5c3aaae7	Merge 'message/messaging_service: guard adding maintenance tenant under cluster feature' from Michał Jadwiszczak In https://github.com/scylladb/scylladb/pull/18729, we introduced a new statement tenant `$maintenance`, but the change wasn't protected by any cluster feature. This wasn't a problem for OSS, since unknown isolation cookie just uses default scheduling group. However, in enterprise that leads to creating a service level on not-upgraded nodes, which may end up in an error if user create maximum number of service levels. This patch adds a cluster feature to guard adding the new tenant. It's done in the way to handle two upgrade scenarios: - version without `$maintenance` tenant -> version with `$maintenance` tenant guarded by a feature - version with `$maintenance` tenant but not guarded by a feature -> version with `$maintenance` tenant guarded by a feature The PR adds `enabled` flag to statement tenants. This way, when the tenant is disabled, it cannot be used to create a connection, but it can be used to accept an incoming connection. The `$maintenance` tenant is added to the config as disabled and it gets enabled once the corresponding feature is enabled. Fixes scylladb/scylladb#20070 Refs scylladb/scylla-enterprise#4403 Closes scylladb/scylladb#19802 * github.com:scylladb/scylladb: message/messaging_service: guard adding maintenance tenant under cluster feature message/messaging_service: add feature_service dependency message/messaging_service: add `enabled` flag to statement tenants	2024-09-17 18:24:34 +02:00
Avi Kivity	1663fbe717	cql3: statement_restrictions: use functional style Instead of a constructor, use a new function analyze_statement_restrictions() as the entry point. It returns an immutable statement_restrictions object. This opens the door to returning a variant, with each arm of the variant corresponding to a different query plan.	2024-09-17 17:13:27 +03:00
Avi Kivity	3169b8e0ec	cql3: statement_restrictions: calculate the index only once find_idx() is called several times. Rename it do_find_idx(), call it just once, store the results, and make find_idx() return the stored results. This simplifies control flow and reduces the risk that successive calls of find_idx return different results.	2024-09-17 17:03:31 +03:00
Avi Kivity	d5c8083b76	cql3: statement_restrictions: make it a const object Make validate_secondary_index_selections() const (it trivially is), and call prepare_indexed_local() / prepared_indexed_global() at the end of the constructor. By making statement_restrictions a const object, reasoning about it can be local (looking at the source file) rather than global (looking at all the interactions of the class with its environment. In fact, we might make it a function one day. Since prepare_indexed_global()/prepare_indexed_local() only mutate _idx_tbl_ck_prefix, which isn't mutated by the rest of the code, the transformation is safe. The corresponding code is removed from select_statement. The removal isn't complete since it still uses some computation, but later deduplication is left for another day.	2024-09-17 17:03:27 +03:00
Sergey Zolotukhin	68740f57c2	cql_server: Add a test for multiple query msg rebounces. The test emulates several LWT(Lightweight Transaction) query rebounces. Currently, the code that processes queries does not expect that a query may be rebounced more than once. It was impossible with the VNodes, but with intruduction of the Tablets, data can be moved between shards by the balancer thus a query can be rebounced to different shards multiple times.	2024-09-17 15:19:56 +02:00
Benny Halevy	65430b9e1b	cql_server::connection: process: rebounce msg if needed Rebounce the msg to another shard if needed, e.g. in the case of tablet migration. An example for that, as given by Tomasz Grabiec: > Bouncing happens when executing LWT statement in > modification_statement::execute_with_condition by returning a > special result message kind. The code assumes that after > jumping to the shard from the bounce request, the result > message is the regular one and not yet another bounce. > There is no problem with vnodes, because shards don't change. > With tablets, they can change at run time on migration. Fixes scylladb/scylladb#15465 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-17 15:09:43 +02:00
Sergey Zolotukhin	f674f522aa	cql_server::connection: process: co-routinize connection::process_on_shard `cql_server::connection::process_on_shard` is made a co-routine to make sure captured objects' lifetime is managed by the source shard, avoiding error prone inter-shard objects transfers.	2024-09-17 14:54:42 +02:00
Nadav Har'El	17deaae463	alternator: make alternator_enforce_authorization live-updateable For no good reason, the "alternator_enforce_authorization" flag (which chooses whether to enable authentication and authorization checks in Alternator) was not live-updatable, so make it so. Both "server" and "executor" objects use this configuration flag, the former is fixed in this patch (to hold a live-updatable reference instead of a copy of a boolean), the latter was already prepared for this change and already held a live-updatable reference. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-17 15:51:16 +03:00
Nadav Har'El	00793059e1	alternator: fix alternator_enforce_authorization=false When the configuration has alternator_enforce_authorization=false, Alternator should not do authentication (check which user signed each request) nor authorization (check if that user has permissions to do each operation). Our implementation forgot to disable the authorization checks when it's configured to false. The (incorrect) assumption was that when alternator_enforce_authorization is configured to false, the CQL 'authenticator' and 'authorizer' configuration is also disabled - so the authorization checks will be no-ops. But we can't assume that: Users are free to configure 'authenticator' and 'authorizer' for use in CQL, and then set alternator_enforce_authorization=false just for Alternator. So this patch adds a new test for this case - when we have authenticator=PasswordAuthenticator, authorizer=CassandraAuthorizer but alternator_enforce_authorization=false, and fixes it to work correctly. The heart of the fix is trivial: the `verify_*_permission()` functions just need to check the alternator_enforce_authorization and return immediately when false. The bigger part of this change is to get the alternator_enforce_authorization into the "executor" object and then to pass it into the verify calls. Although alternator_enforce_authorization is not YET live updatable, this code is prepared for the future that it may become live updatable, so the executor object saves not the boolean value of this flag, but a live-updatable reference to it. Fixes #20619 Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-17 15:50:00 +03:00
Nadav Har'El	76af7c0389	alternator: improve error message when unauthenticated When access-control checks report permission denied, we want to report the name of the authenticated role (the role signing the request) which didn't have the permission. When authentication was disabled, and there is no authenticated role, we printed the fake name "anonymous", but this can confuse users (it confused me!) to think there's an actual role named "anonymous". So let's change that string to "<anonymous>" with angle brackets - it makes it more obvious that this isn't a real role, but actually an anonymous request. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-17 15:44:29 +03:00
Tomasz Grabiec	e70ce4d6ed	gdb: Introduce "scylla sstable-dump-cached-index" command	2024-09-17 14:41:18 +02:00
Tomasz Grabiec	9f0eed263d	gdb: Introduce "scylla sstable-promoted-index" command	2024-09-17 14:41:13 +02:00
Nadav Har'El	3543bf14e9	alternator: avoid use-after-free in RBAC While auditing the code, I noticed that the current Alternator access control checks have code like: ``` return client_state.check_has_permission(auth::command_desc( permission_to_check, auth::make_data_resource(schema->ks_name(), schema->cf_name()))).then( ``` There's a problem here - it turns out that, unfortunately, command_desc holds a reference to the "resource" object - not a copy. So the temporary object returned by make_data_resource may be freed and then used... Curiously, we've not seen a bug caused by this in practice (not even in debug build mode), but better safe than sorry, so this patch changes the code in one of two ways: 1. Code using coroutines can keep the "resource" as a variable on the stack. 2. Code using continuations needs to hold the "resource" with do_with(), but since this already incurs the cost of an extra allocation (even in the successful case), might as well just switch to using coroutines and have less ugly code. This patch does not change any functionality, and all the tests seem to work before and after it the same. Signed-off-by: Nadav Har'El <nyh@scylladb.com> hello	2024-09-17 15:41:09 +03:00
Tomasz Grabiec	2c463ead59	gdb: Fix range printer for singular ranges Before, it printed [x, +inf) instead of {x}	2024-09-17 14:30:28 +02:00
Andrei Chekun	bbb6c3c2ff	test.py: Add resource consumption metrics This PR adds the possibility to gather resource consumption metrics. The collected metrics can be used to compare performance before and after specific changes aimed at increasing performance. Currently, this functionality works only in manual mode, and this is just raw data. Later on, these metrics can be used in Jupyter notebook to analyze and visualize how the resources are used and can provide the insight on how to improve it. This PR is a first insight after gathering these metrics. Add the possibility to gather resource consumption for the test.py execution. SQLite DB will be created with different performance metrics that will allow comparing the resource consumption between changes. The DB will be in the tmp directory that by default set to testlog. Across the runs, the DB will not be deleted, so each new run will just add information to the existing DB. Parameter --get-metrics was added to switch on or off the metrics gathering. By default, it's switched on. Closes: scylladb/qa-tasks#1666 Closes: scylladb/qa-tasks#1707 Closes scylladb/scylladb#19881	2024-09-17 15:22:34 +03:00
Benny Halevy	39ce358d82	time_window_compaction_strategy: get_reshaping_job: restrict sort of multi_window vector to its size Currently the function calls boost::partial_sort with a middle iterator that might be out of bound and cause undefined behavior. Check the vector size, and do a partial sort only if its longer than `max_sstables`, otherwise sort the whole vector. Fixes scylladb/scylladb#20608 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#20609	2024-09-17 15:05:37 +03:00
Tomasz Grabiec	adf99402c5	Merge 'readers/flat_mutation_reader_v2: call set_close_required() from consume()' from Botond Dénes The `consume()` variants just forward the call to the `_impl` method with the same name. The latter, being a member of `::impl`, will bypass the top level `fill_buffer()`, etc. methods and thus will never call `set_close_required()`. Do this in the top-level `consume()` methods instead, to ensure a reader, on which only `consume()` is called, and then is destroyed, will complain as it should (and abort). Only one place was found in core code, which didn't close the reader: `split_mutation() in `mutation/mutation.cc` and this reader is the "from-mutation" one which has no real close routine. All other places were in tests. All this is to say, there were no real bugs uncovered by this PR. Fixes #16520 Improvement, no backport required. Closes scylladb/scylladb#16522 * github.com:scylladb/scylladb: readers/flat_mutation_reader_v2: call set_close_required() from consume*() test/boost/sstable_compaction_test: close reader after use test/boost/repair_test: close reader after use mutation/mutation: split_mutation(): close reader after use	2024-09-17 13:21:34 +02:00
Anna Mikhlin	66c0814c33	Update ScyllaDB version to: 6.3.0-dev	2024-09-17 13:43:04 +03:00
Botond Dénes	6250ff18eb	Merge 'sstable: s/crawling_sstable_mutation_reader/sstable_full_scan_reader' from Kefu Chai "crawling" is a little bit obscure in this context. so let's rename this class to reflect the fact that this reader only reads the entire content of the sstable. both crawling reader for kl and mx formats are renamed. also, in order to be consistent, all "crawling reader" in variable names are updated as well. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#20599 * github.com:scylladb/scylladb: sstable: s/crawling_sstable_mutation_reader/sstable_full_scan_reader sstable/mx/reader: add comment for mx_crawling_sstable_mutation_reader	2024-09-17 11:55:08 +03:00
Pavel Emelyanov	ebfa73e004	s3/client: Don't move file from write_body's lambda Requests sent by S3 are retriable, so when request.write_body() is called, it should keep everything intact in case http client will call it again. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20579	2024-09-17 09:48:09 +03:00
Tzach Livyatan	cb864b11d8	Update client-node-encryption: OpsnSSL is FIPS enabled Closes scylladb/scylladb#19705	2024-09-17 09:47:07 +03:00
Botond Dénes	f32e67cb9e	Merge 'Make sstables without on-disk path' from Pavel Emelyanov New sstables for a table are created by the table::make_sstable() method. The method then calls sstables_manager::make_sstable() and passes there a path to component files which, in turn, sits on table::config. Since some time ago having an on-disk path for an sstable had become optional, as sstables could be put on S3 storage without local paths involved. In that case the aforementioned "path" is ~~ab~~used as a key in the system.sstables registry, that references a record with information used to retrieve URLs of sstables' objects. This PR removes the "path" argument from sstables_manager::make_sstable() and its sstable_sdirectory peer. The details of sstables' location are moved onto storage_options and depend on storage type. For now in both storage types this location is still the good-old $datadir/$keyspace/$table-$uuid string. S3 storage needs to be patched more to use more elegant "location" value. Eventually the `table::config::{datadir\|all_datadirs}` will be removed, this PR is the step towards it. closes: #12707 Closes scylladb/scylladb#20542 * github.com:scylladb/scylladb: table: Use storage options to clean the storage sstables/storage: Re-use ocally generated vector of paths sstables/storage: Visit options once to initialize storage sstables_manager: Return table storage options when initalizing storage sstables/storage: Fix indentation after previous patch table: Move datadirs initialization parallelism to storage level sstables/storage: Split the visitor's overloaded functor restore: Don't use table_dir to construct sstable_directory sstable_directory: Remove table_dir field sstable_directory: Use options details in lister sstables_manager: Remove table_dir from make_sstable() sstables: Remove table_dir from sstable constructor sstables/storage: Remove sstring dir from make_storage() sstables/storage: Use options to construct tests: Properly initialize storage options with "dir" distributed_loader: Create S3 options with prefix for restore storage_options: Add special-purpose local options maker storage_options: Keep local path / s3 prefix onboard table: Get another options when initializing storage	2024-09-17 09:41:21 +03:00
Botond Dénes	a4a8cad97f	Merge 'atomic_delete: allow deletion of sstables from several prefixes' from Benny Halevy Allow create_pending_deletion_log to delete a bunch of sstables potentially resides in different prefixes (e.g. in the base directory and under staging/). The motivation arises from table::cleanup_tablet that calls compaction_group::cleanup on all cg:s via cleanup_compaction_groups. Cleanup, in turn, calls delete_sstables_atomically on all sstables in the compaction_group, in all states, including the normal state as well as staging - hence the requirement to support deleting sstables in different sub-directories. Also, apparently truncate calls delete_atomically for all sstables too, via table::discard_sstables, so if it happened to be executed during view update generation, i.e. when there are sstables in staging, it should hit the assertion failure reported in https://github.com/scylladb/scylladb/issues/18862 as well (although I haven't seen it yet, but I see no reason why it would happen). So the issue was apparently present since the initial implementation of the pending_delete_log. It's just that with tablet migration it is more likely to be hit. Fixes scylladb/scylladb#18862 Needs backport to 6.0 since tablets require this capability Closes scylladb/scylladb#19555 * github.com:scylladb/scylladb: sstable_directory: create_pending_deletion_log: place pending_delete log under the base directory sstables: storage: keep base directory in base class sstables: storage: define opened_directory in header file sstable_directory: use only dirlog	2024-09-17 08:30:40 +03:00
Kefu Chai	df7f332a58	sstable: s/crawling_sstable_mutation_reader/sstable_full_scan_reader "crawling" is a little bit obscure in this context. so let's rename this class to reflect the fact that this reader only reads the entire content of the sstable. both crawling reader for kl and mx formats are renamed. also, in order to be consistent, all "crawling reader" in variable names are updated as well. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-17 10:39:37 +08:00
Kefu Chai	c1ed2f0ea4	sstable/mx/reader: add comment for mx_crawling_sstable_mutation_reader to explain its typical usage. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-17 10:39:25 +08:00
Lakshmi Narayanan Sreethar	626f55a2ea	compaction: run cleanup under maintenance scheduling group The cleanup compaction task is a maintenance operation that runs after topology changes. So, run it under the maintenance scheduling group to avoid interference with regular compaction tasks. Also remove the share allocations done by the cleanup task, as they are unnecessary when running under the maintenance group. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com> Closes scylladb/scylladb#20582	2024-09-16 16:58:43 +03:00
Michał Jadwiszczak	b4b91ca364	message/messaging_service: guard adding maintenance tenant under cluster feature Set `enabled` flag for `$maintenance` tenant to false and enable it when `MAINTENANCE_TENANT` feature is enabled.	2024-09-16 15:34:36 +02:00
Michał Jadwiszczak	71a03ef6b0	message/messaging_service: add feature_service dependency	2024-09-16 15:33:40 +02:00
Michał Jadwiszczak	d44844241d	message/messaging_service: add `enabled` flag to statement tenants Adding a new tenant needs to be done under cluster feature protection. However it wasn't the case for adding `$maintenance` statement tenant and to fix it we need to support an upgrade from node which doesn't know about maintenance tenant at all and from one which uses it without any cluster feature protection. This commit adds `enabled` flag to statement tenants. This way, when the tenant is disabled, it cannot be used to create a connection, but it can be used to accept an incoming connection.	2024-09-16 15:31:04 +02:00
Michał Jadwiszczak	de7acbad8b	test/cql-pytest: add test for `SELECT ... USING SERVICE LEVEL`	2024-09-16 14:31:43 +02:00
Michał Jadwiszczak	8255c61f5f	cql3/Cql.g: extend grammar to allow `SELECT ... USING SERVICE LEVEL`	2024-09-16 14:31:32 +02:00
Michał Jadwiszczak	af6dc78025	cql3/statements/select_statement: use service level timeout Use service level timeout in selecte statement when specified. `USING TIMEOUT` have higher priority in timeout definition.	2024-09-16 13:48:48 +02:00
Michał Jadwiszczak	2e545c915b	cql3/attributes: add service level name field In next patches, we will allow to do `SELECT ... USING SERVICE LEVEL sl_name`. To do it, we need to extend `cql3::attributes` with service level name.	2024-09-16 13:48:43 +02:00
Michał Jadwiszczak	b9b326c2bb	qos/service_level_controller: add method to check if service level exists in cache There is `service_level_controller::get_service_level()` method, which searches for service level in the controller cache and returns default service level if SL with given name doesn't exist. Added method allows to check whether a service level exists in the controller cache.	2024-09-16 12:41:15 +02:00
Pavel Emelyanov	bf5021e735	test: Remove sstables::test::binary_search() That's the most mysterious wrapper in this set as it doesn't need sstable itself at all, it just duplicates the existing non-class function out there. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-16 12:51:35 +03:00
Pavel Emelyanov	309d315af7	test: Remove sstables::test::move_summary() This one is a bit tricky, as it needs to modify the sstables's summary. However, the sstables::test::_summary() one returns mutable reference and the only caller can use it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-16 12:50:48 +03:00
Pavel Emelyanov	deec952111	test: Remove sstables::test::read_toc() The sstable::read_toc() is public method, use it directly. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-16 12:50:19 +03:00
Pavel Emelyanov	25cd8ccdd8	test: Remove sstables::test::get_summary() Same as previous patch -- callers can come with const reference to summary, so they can live with existing public sstable::get_summary(). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-16 12:49:39 +03:00
Pavel Emelyanov	f714ac9b48	test: Remove sstables::test::get_statistics() Just call the public sstable::get_statistics(). The callers would get const reference on it, but they don't need more than that. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-16 12:48:43 +03:00
Pavel Emelyanov	53afa583e8	test: Remove sstables::test::data_read() The wrapper just changes the order of arguments for a public method. Drop it, and call the wrapee directly. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-16 12:47:59 +03:00
Avi Kivity	e4cab3a5e9	cql3: statement_restrictions: drop accessors for single-column key restrictions No longer used.	2024-09-16 12:15:14 +03:00
Avi Kivity	626acf416e	cql3: selection: adjust indentation	2024-09-16 12:15:14 +03:00
Avi Kivity	c443d922ea	cql3: selection: delete empty loop Our refactoring left a loop with no body, delete it.	2024-09-16 12:15:14 +03:00
Avi Kivity	56e8a4c931	cql3: statement_restrictions, selection: fold multi-column restrictions into row-level filter When filtering, we apply single-column and multi-column filters separately. This is completely unnecessary. Find the multi-column filters during prepare time and append them to the row-level filter. This slightly changes the original: in the original, if we had a multi-column filter, we applied all of the restrictions. But hopefully if we check for multi-column filters, that's what we need.	2024-09-16 12:15:14 +03:00
Avi Kivity	a6d81806c0	cql3: statement_restrictions, selection: merge clustering key filter and regular columns filter The two filters are used in the same way: check the filter, return false if it matches. Unify the two filters into a clustering_row_level_filter. Since one of the two filters wasn't std::optional, we take the liberty of making the combined filter non-optional.	2024-09-16 12:15:03 +03:00
Avi Kivity	2933a2f118	cql3: statement_restrictions, selection: merge partition key filter and static columns filter The two filters are used in the same way: check the filter, set a boolean flag if it matches, return false. The two boolean flags are in turn checked in the same way. Unify the two filters into a partition_level_filter. Since one of the two filters wasn't std::optional, we take the liberty of making the combined filter non-optional.	2024-09-16 12:10:49 +03:00
Avi Kivity	870d1c16f7	scripts: fix bin/cqlsh shortcut Since `3c7af28725`, the cqlsh submodule no longer contains a bin/cqlsh shell script. This broke the supermodule's bin/cqlsh shortcut. Fix it by invoking cqlsh.py directly. Closes scylladb/scylladb#20591	2024-09-16 09:52:29 +03:00
Botond Dénes	ea29fe579b	Merge 'replica: ignore cleanup of deallocated storage group' from Aleksandra Martyniuk Cleanup of a deallocated tablet throws an exception. Since failed cleanup is retried, we end up in an infinite loop. Ignore cleanup of deallocated storage groups. Fixes: #19752. Needs to be backported to all branches with tablets (6.0 and later) Closes scylladb/scylladb#20584 * github.com:scylladb/scylladb: test: check if cleanup of deallocated sg is ignored replica: ignore cleanup of deallocated storage group	2024-09-16 09:22:56 +03:00
Gleb Natapov	695f112795	paxos_state: release semaphore units before checking if a semaphore can be dropped To drop a semaphore it should not be held by anyone, so we need to release out units before checking if a semaphore can be dropped. Fixes: scylladb/scylladb#20602 Closes scylladb/scylladb#20607	2024-09-15 21:21:03 +03:00
Kefu Chai	028410ba58	mutation_writer: use bucket parameter instead of using it->first as `_bucket` is an `unordered_map<bucket_id, timestamp_bucket_writer>`, when writing to a given bucket, we try to create a writer with the specified bucket id, so the returned iterator should point to a node whose `first` element is always the bucket id. so, there is no need to reference `it` for the bucket id, let's just reference the parameter. simpler this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20598	2024-09-15 20:05:12 +03:00
Kefu Chai	49f232f405	compaction: fix a typo in comment s/expection/exception/ Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20594	2024-09-15 16:09:01 +03:00
Avi Kivity	807153a9ed	cql3: selection: filter regular and static rows as a single expression each Instead of filtering regular and static columns column by column, call is_satisfied_by() for an expression containing all the static columns predicates, and one for all the regular column. We cannot have one expression, since the code sets _current_static_row_does_not_match only for static columns. Note the fix for #20485 is now implicit, since the evaluation machinery will treat missing regular columns as NULL.	2024-09-15 14:33:57 +03:00
Avi Kivity	3c71096479	cql3: statement_restrictions: collect regular column and static column filters into single expressions Similar to previous work with clustering and partition key, expose static and reglar column filters as single expressions. Since we don't currently expose a boolean for whether those filters exist, we expose them now as non-optionals. In any case evaluating an empty conjunction is plenty fast.	2024-09-15 14:33:57 +03:00
Avi Kivity	ec2898afe9	cql3: selection: filter clustering key as a single expression Instead of filtering the clustering key column by column, call is_satisfied_by() for an expression containing all the clustering key predicates. The check for clustering_key.empty() is removed; the evaluation machinery is able to handle partial clustering keys. In fact if we add IS NULL, we have to evaluate as an empty clustering key should match.	2024-09-15 14:33:57 +03:00
Avi Kivity	318d653d80	cql3: statement_restrictions: expose filter for clustering key cql3::selection performs filtering by consulting ck_restrictions_need_filtering() and get_single_column_clustering_key_restrictions() (which is a map of column definition to expressions). Make them available in one nice package as an optional<expression>. When the optional is engaged, filtering is needed, and the expression in the equivalent of all of the map.	2024-09-15 14:33:57 +03:00
Avi Kivity	0bd2f12922	cql3: selection: filter partition key as a single expression Instead of filtering the partition key column by column, call is_satisfied_by() for an expression containing all the partition key predicates.	2024-09-15 14:33:56 +03:00
Avi Kivity	21cb91077f	cql3: statement_restrictions: expose filter for partition key cql3::selection performs filtering by consulting pk_restrictions_need_filtering() and get_single_column_partition_key_restrictions() (which is a map of column definition to expressions). Make them available in one nice package as an optional<expression>. When the optional is engaged, filtering is needed, and the expression in the equivalent of all of the map.	2024-09-15 14:33:56 +03:00
Avi Kivity	a453221314	cql3: statement_restrictions: remove relations used for indexing from filtering statement_restrictions does not name columns that were used for a secondary index for selection for filtering, since accessing the index "pre-filters" these columns. However, it keeps the relations that contain these columns. This makes it impossible (besides unnecessary) to evaluate the relations, as the columns they reference aren't selected. The reason this works now is that result_set_builder::restrictions_filter::do_filter() iterates on selected columns, matching them to relations, then execute the matched relation. A relation that references an unselected column is invisible to do_filter(). We wish to filter using complete expressions, rather than fragments, so as a first step remove these unnecessary and unusable relations while we choose which columns are necessary for filtering. calculate_column_defs_for_filtering is renamed to remind us of the extra work done.	2024-09-15 14:33:56 +03:00
Avi Kivity	ba8c2014bf	cql3: statement_restrictions: bail out of find_idx if !_uses_secondary_index The condition seems trivial, but wasn't implemented, without ill effects so far. With the following patches, calculate_column_defs_for_filtering() becomes confused as it selects an indexing code path even when !_uses_secondary_index, triggered by the reproducer of #10300.	2024-09-15 14:33:56 +03:00
Avi Kivity	65ba19323c	cql3: statement_restrictions, modification_statement: pass correct value of check_indexes Our UPDATE/INSERT/DELETE statements require a full primary/partition key and therefore never use indexes; fix the check_index parameter passed from modification_statement. So far the bug is benign as we did not take any action on the value. Make the parameter non-default to avoid such confusion in the future.	2024-09-15 14:33:56 +03:00
Avi Kivity	71ea3200ba	cql3: statement_restrictions: correct mismatched clustering/partition restrictions references The second loop of calculate_column_defs_for_filtering() finds clustering keys that are used for filtering, minus and clustering keys that happen to be used for secondary indexing. However, to check whether the clustering key is used for secondary indexing, it looks up in _single_column_partition_key_restrictions, which contains partition key restrictions. The end result is that we select a column which ends the partition key for the secondary index, and so is unnecessary. We do a little more work, but the bug is benign. Nevertheless, fix it, as it interferes with following work.	2024-09-15 14:33:56 +03:00
Avi Kivity	33db14e7d5	cql3: statement_restrictions: precalculate get_column_defs_for_filtering() get_column_defs_for_filtering() names all the columns that are required for filtering. While doing that, it skips over columns that are participate in indexing (primary or secondary), since the index "pre-filters" the query. We wish to make use of this skipping. As a first step, call the calculation from the constructor, so we have control over when it is executed.	2024-09-15 14:33:56 +03:00
Avi Kivity	251ad4fcd0	cql3: selection: do_filter(): push static/regular row glue to higher level Currently, for each column we call get_non_pk_values() to transform the way we get the information (query::result_row_view) to the way the expression evaluation machinery wants it (vector<managed_bytes_opt>). Call it just once outside the loop.	2024-09-15 14:33:56 +03:00
Avi Kivity	b9bc783418	cql3: selection: don't ignore regular column restriction if a regular row is not present If a regular row isn't present, no regular column restriction (say, r=3) can pass since all regular columns are presented as NULL, and we don't have an IS NULL predicate. Yet we just ignore it. Handle the restriction on a missing column by return false, signifying the row was filtered out. We have to move the check after the conditional checking whether there's any restriction at all, otherwise we exit early with a false failure. Unit test marked xfail on this issue are now unmarked. A subtest of test_tombstone_limit is adjusted since it depended on this bug. It tested a regular column which wasn't there, and this bug caused the filter to be ignored. Change to test a static column that is there. A test for a bug found while developing the patch is also added. It is also tested by test_tombstone_limit, but better to have a dedicated test. Fixes #10357 Closes scylladb/scylladb#20486	2024-09-15 13:44:16 +03:00
Botond Dénes	6d8e9645ce	test/*/run: restore --vnodes into working order This option was silently broken when --enable-tablet's default changed from false to true. The reason is that when --vnodes is passed, run only removes --enable-tablets=true from scylla's command line. With the new default this is not enough, we need to explicitely disable tablets to override the default. Closes scylladb/scylladb#20462	2024-09-13 17:10:09 +03:00
Pavel Emelyanov	f850681b14	table: Use storage options to clean the storage Like it was done for table::init_storage(), patch the table::destroy_storage() not to mess with datadir path and rely on storage options only. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	3aea7bebb7	sstables/storage: Re-use ocally generated vector of paths A cleanup after prefious patch -- in order to create storage options for table the local initialization code can re-use the vector of paths that it hag generated in the same call to create table directory layout. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	7c34724509	sstables/storage: Visit options once to initialize storage The init_table_storage() method now does it twice -- one time to initialize the storage, another one to create new options for table. Both can be merged, thus making table storage options initialization better encapsulated for local/s3 cases. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	311fb906be	sstables_manager: Return table storage options when initalizing storage Now the table::init_storage() calls sstables manager two times -- first, to get storage options, second, to initialize the storage with obtained options. Merge two calls into one. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	918ec00c1d	sstables/storage: Fix indentation after previous patch Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	f1e4367439	table: Move datadirs initialization parallelism to storage level The table::init_table_storage() calls sstables_manager's storage initialization for each of the datadirs found on config. That's not great, it's sstables manager (and its storage) that know if table needs to mess with datadirs or not. This patch moves the loop to storage.cc. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	b6b3a477c5	sstables/storage: Split the visitor's overloaded functor The main goal is to have init_table_storage() overload for local options as standalone function. This makes next patching simpler. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	30c8d89f97	restore: Don't use table_dir to construct sstable_directory Continuation of the previous patch patching the special-purpose sstable directory constructor that's used by restore-from-s3-backup code. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	af14408052	sstable_directory: Remove table_dir field It's no longer needed -- both, lister and making sstable, work with having storage options at hand. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	f403728aa4	sstable_directory: Use options details in lister This class is very similar to sstables::storage one -- it also needs path or s3 prefix to construct. Now when this information is stored on storage_options, it's better to stick to it, not to the argument. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	36863d4ad0	sstables_manager: Remove table_dir from make_sstable() It used to be passed to sstable constructor, but now it doesn't need this argument. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	0764eca553	sstables: Remove table_dir from sstable constructor It used to be passed to storage constructor, now storage works with options only and this argument is no longer needed. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	d79ae1f02b	sstables/storage: Remove sstring dir from make_storage() Now the directory/s3 prefix is propagated via storage options. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	65a19df8ef	sstables/storage: Use options to construct All callers of make_sstable are now patched to provide correct storage options with path/prefix set. The make_storage() helper can switch to using it. Respectively, it's good to make sure that the storage is created with table options that have path/prefix. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	4425cf54c6	tests: Properly initialize storage options with "dir" Most of the tests work with local storage options. Some support S3 options as well. Whatever it is, when creating an sstable, tests need to put proper "dir" on the options, this patch does so. In fact, storage options for tests are created together with the test-env, and ideally this is the place where dir should be assigned on it. However, there are still places that explicitly specify path they want to see sstables at, for those the new temporary options should be constructed. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	33bc9e7112	distributed_loader: Create S3 options with prefix for restore Restore-from-backup code wants to collect sstables from remote S3. For that it constructs S3 options, and now it needs to put prefix on it as well. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:32:39 +03:00
Pavel Emelyanov	56111a50cd	storage_options: Add special-purpose local options maker Lost of code (in tools and tests) explicitly deal with local sstables and need to create options for it. Currently default-constructing options generates local ones, but without the directory path. Add a helper that creates local options with path and patch callers. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:32:39 +03:00
Pavel Emelyanov	95e60cde9f	storage_options: Keep local path / s3 prefix onboard Now when tables keep their own copy of storage options, it's possible for each table to add table-specific information on it. Namely -- path for local storage and prefix for S3 one (in fact, it's not a "prefix", but a key in sstables registry, but fixing it is beyond the scope of this set). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:32:32 +03:00
Pavel Emelyanov	14976fda73	table: Get another options when initializing storage Right now the table's storage_options life starts in cql, and shortly after the lw-shared-pointer to options is put on keyspace metadata. Later, when the table is created the pointer from keyspace is copied on the table via its contructor. Next patches will extend the options pointed to by a table, and the extension is going to be different for different tables. For that, each table needs to have its private options and this patch prepares for that. For now table directly calls sstables/storage code to get the options from, but it's temporary, soon the options will be created via sstables manager together with initialising the storage itself. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:32:32 +03:00
Nadav Har'El	f255391d52	cql-pytest: translate Cassandra's tests for arithmetic operators This is a translation of Cassandra's CQL unit test source file OperationFctsTest.java into our cql-pytest framework. This is a massive test suite (over 800 lines of code) for Cassandra's "arithmetic operators" CQL feature (CASSANDRA-11935), which was added to Cassandra almost 8 years ago (and reached Cassandra 4.0), but we never implemented it in Scylla. All of the tests in suite fail in ScyllaDB due to our lack of this feature: Refs #2693: Support arithmetic operators One test also discovered a new issue: Refs #20501: timestamp column doesn't allow "UTC" in string format All the tests pass on Cassandra. Some of the tests insist on specific error message strings and specific precision for decimal arithmetic operations - where we may not necessarily want to be 100% compatible with Cassandra in our eventual implementation. But at least the test will allow us to make deliberate - and not accidental - deviations from compatibility with Cassandra. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20502	2024-09-13 14:52:59 +03:00
Botond Dénes	d3a9654fcc	Merge 'Make use of async() context in sstable_mutation_test' from Pavel Emelyanov This test runs all its cases in seastar thread, but still uses .then() continuations in some of them. This PR converts all continuations into plain .get()-s. Closes scylladb/scylladb#20457 * github.com:scylladb/scylladb: test: Restore indentation after previous changes test: Threadify tombstone_in_tombstone2() test: Threadify range_tombstone_reading() test: Threadify tombstone_in_tombstone() test: Threadify broken_ranges_collection() test: Threadify compact_storage_dense_read() test: Threadify compact_storage_simple_dense_read() test: Threadify compact_storage_sparse_read() test: Simplify test_range_reads() counting test: Simplify test_range_reads() inner loop test: Threadify test_range_reads() itself test: Threadify test_range_reads() callers test: Threadify generate_clustered() itself test: Threadify generate_clustered() callers test: Threadify test_no_clustered test test: Threadify nonexistent_key test	2024-09-13 14:09:53 +03:00
Aleksandra Martyniuk	2c4b1d6b45	test: check if cleanup of deallocated sg is ignored	2024-09-13 13:00:58 +02:00
Aleksandra Martyniuk	20d6cf55f2	replica: ignore cleanup of deallocated storage group Currently, attempt to cleanup deallocated storage group throws an exception. Failed tablet cleanup is retried, stucking in an endless loop. Ignore cleanup of deallocated storage group.	2024-09-13 13:00:53 +02:00
Botond Dénes	cb30271d29	readers/flat_mutation_reader_v2: call set_close_required() from consume() The `consume()` variants just forward the call to the `_impl` method with the same name. The latter, being a member of `::impl`, will bypass the top level `fill_buffer()`, etc. methods and thus will never call `set_close_required()`. Do this in the top-level `consume()` methods instead, to ensure a reader, on which only `consume()` is called, and then is destroyed, will complain as it should (and abort). operator()() was also missing `set_close_required()`, fix that too.	2024-09-13 06:52:26 -04:00
Botond Dénes	fbed280cd5	test/boost/sstable_compaction_test: close reader after use	2024-09-13 06:52:26 -04:00
Botond Dénes	116b044fec	test/boost/repair_test: close reader after use	2024-09-13 06:52:26 -04:00
Botond Dénes	1a11f9cf95	mutation/mutation: split_mutation(): close reader after use	2024-09-13 06:52:26 -04:00
Andrei Chekun	bad7407718	test.py: Add support for BOOST_DATA_TEST_CASE Currently, test.py will throw an error if the test will use BOOST_DATA_TEST_CASE. test.py as a first step getting all test functions in the file, but when BOOST_DATA_TEST_CASE will be used the output will have additional lines indicating parametrized test that test.py can not handle. This commit adds handling this case, as a caveat all tests should start from 'test' or they will be ignored. Closes: #20530 Closes scylladb/scylladb#20556	2024-09-13 13:44:26 +03:00
Botond Dénes	7cb8cab2ae	Merge 'Remove make_shared_schema() helper' from Pavel Emelyanov This function was obsoleted by schema_builder some time ago. Not to patch all its callers, that helper became wrapper around it. Remained users are all in tests, and patching the to use builder directory makes the code shorter in many cases. Closes scylladb/scylladb#20466 * github.com:scylladb/scylladb: schema: Ditch make_shared_schema() helper test: Tune up indentation in uncompressed_schema() test: Make tests use schema_builder instead of make_shared_schema	2024-09-13 12:25:10 +03:00
Pavel Emelyanov	730731da4a	test: Remove unused table config from max_ongoing_compaction_test The local config is unused since #15909, when the table creation was changed to use env's facilities. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20511	2024-09-13 12:21:56 +03:00
Pavel Emelyanov	4c77f474ed	test: Remove unused upload_path local variable Since #14152 creation of an sstable takes table dir and its state. The test in question wants to create and sstable in upload/ subdir and for that it used to maintain full "cf.dir/upload" path, which is not required any more. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20514	2024-09-13 12:21:00 +03:00
Pavel Emelyanov	e9a1c0716f	test: Use sstables::test_env to make sstables for directory test This is continuation of #20431 in another test. After #20395 it's also possible to remove unused local dir variables. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20541	2024-09-13 12:19:59 +03:00
Botond Dénes	4fb194117e	Merge 'Generalize multipart upload implementations in S3 client' from Pavel Emelyanov There are two currently -- upload_sink_base and do_upload_file. This PR merges as much code as possible (spoiler: it's already mostly copy-n-pase-d, so squashing is pretty straightforward) Closes scylladb/scylladb#20568 * github.com:scylladb/scylladb: s3/client: Reuse class multipart_upload in do_upload_file s3/client: Split upload_sink_base class into two	2024-09-13 10:35:10 +03:00
Kefu Chai	cf1f90fe0c	auth: remove unused #include the `seastar/core/print.hh` header is no longer required by `auth/resource.hh`. this was identified by clang-include-cleaner. As the code is audited, wecan safely remove the #include directive. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20575	2024-09-13 09:49:05 +03:00
Botond Dénes	c7c5817808	Merge 'Improve timestamp heuristics for tombstone garbage collection' from Benny Halevy When purging regular tombstone consult the min_live_timestamp, if available. This is safe since we don't need to protect dead data from resurrection, as it is already dead. For shadowable_tombstones, consult the min_memtable_live_row_marker_timestamp, if available, otherwise fallback to the min_live_timestamp. If we see in a view table a shadowable tombstone with time T, then in any row where the row marker's timestamp is higher than T the shadowable tombstone is completely ignored and it doesn't hide any data in any column, so the shadowable tombstone can be safely purged without any effect or risk resurrecting any deleted data. In other words, rows which might cause problems for purging a shadowable tombstone with time T are rows with row markers older or equal T. So to know if a whole sstable can cause problems for shadowable tombstone of time T, we need to check if the sstable's oldest row marker (and not oldest column) is older or equal T. And the same check applies similarly to the memtable. If both extended timestamp statistics are missing, fallback to the legacy (and inaccurate) min_timestamp. Fixes scylladb/scylladb#20423 Fixes scylladb/scylladb#20424 > [!NOTE] > no backport needed at this time > We may consider backport later on after given some soak time in master/enterprise > since we do see tombstone accumulation in the field under some materialized views workloads Closes scylladb/scylladb#20446 * github.com:scylladb/scylladb: cql-pytest: add test_compaction_tombstone_gc sstable_compaction_test: add mv_tombstone_purge_test sstable_compaction_test: tombstone_purge_test: test that old deleted data do not inhibit tombstone garbage collection sstable_compaction_test: tombstone_purge_test: add testlog debugging sstable_compaction_test: tombstone_purge_test: make_expiring: use next_timestamp sstable, compaction: add debug logging for extended min timestamp stats compaction: get_max_purgeable_timestamp: use memtable and sstable extended timestamp stats compaction: define max_purgeable_fn tombstone: can_gc_fn: move declaration to compaction_garbage_collector.hh sstables: scylla_metadata: add ext_timestamp_stats compaction_group, storage_group, table_state: add extended timestamp stats getters sstables, memtable: track live timestamps memtable_encoding_stats_collector: update row_marker: do nothing if missing	2024-09-13 08:56:51 +03:00
Takuya ASADA	3cd2a61736	dist: drop scylla-jmx Since JMX server is deprecated, drop them from submodule, build system and package definition. Related scylladb/scylla-tools-java#370 Related #14856 Signed-off-by: Takuya ASADA <syuu@scylladb.com> Closes scylladb/scylladb#17969	2024-09-13 07:59:45 +03:00
Botond Dénes	fc9804ec31	Update tools/java submodule * tools/java 0b4accdd...e505a6d3 (1): > [C-S] Make it use DCAwareRoundRobinPolicy unless rack is provided Closes scylladb/scylladb#20562	2024-09-13 06:30:04 +03:00
Takuya ASADA	0ac450de05	scylla_raid_setup: configure SELinux file context On RHEL9, systemd-coredump fails to coredump on /var/lib/scylla/coredump because the service only have write acess with systemd_coredump_var_lib_t. To make it writable, we need to add file context rule for /var/lib/scylla/coredump, and run restorecon on /var/lib/scylla. Fixes #20573	2024-09-13 04:31:52 +09:00
Takuya ASADA	56c971373c	scylla_coredump_setup: fix SELinux configuration for RHEL9 Seems like specific version of systemd pacakge on RHEL9 has a bug on SELinux configuration, it introduced "systemd-container-coredump" module to provide rule for systemd-coredump, but not enabled by default. We have to manually load it, otherwise it causes permission error. Fixes #19325	2024-09-13 04:31:16 +09:00
Pavel Emelyanov	17e7d3145c	s3/client: Reuse class multipart_upload in do_upload_file Uploading a file is implemented by the do_upload_file class. This class re-implements a big portion of what's currently in multipart_upload one. This patch makes the former class inherit from the latter and removes all the duplication from it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-12 18:38:16 +03:00
Pavel Emelyanov	14b741afc9	s3/client: Split upload_sink_base class into two This class implements two facilities -- multipart upload protocol itself plus some common parts of upload_sink_impl (in fact -- only close() and plugs put(packet)). This patch aplits those two facilities into two classes. One of them will be re-used later. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-12 18:00:19 +03:00
Sergey Zolotukhin	612a141660	raft: Fix race condition on override_snapshot_thresholds. When the server_impl::applier_fiber is paused by a co_await at line raft/server.cc:1375: ``` co_await override_snapshot_thresholds(); ``` a new snapshot may be applied, which updates the actual values of the log's last applied and snapshot indexes. As a result, the new snapshot index could become higher than the old value stored in _applied_idx at line raft/server.cc:1365, leading to an assertion failure in log::last_conf_for(). Since error injection is disabled in release builds, this issue does not affect production releases. This issue was introduced in the following commit `9dfa041fe1`, when error injection was added to override the log snapshot configuration parameters. How to reproduce: 1. Build debug version of randomized_nemesis_test ``` ninja-build build/debug/test/raft/randomized_nemesis_test ``` 2. Run ``` parallel --halt now,fail=1 -j20 'build/debug/test/raft/randomized_nemesis_test \ --run_test=test_frequent_snapshotting -- -c2 -m2G --overprovisioned --unsafe-bypass-fsync 1 \ --kernel-page-cache 1 --blocked-reactor-notify-ms 2000000 --default-log-level \ trace > tmp/logs/eraseme_{}.log 2>&1 && rm tmp/logs/eraseme_{}.log' ::: {1..1000} ``` Fixes scylladb/scylladb#20363 Closes scylladb/scylladb#20555	2024-09-12 16:19:27 +02:00
Aleksandra Martyniuk	59fba9016f	docs: operating-scylla: add task manager docs Admin-facing documentation of task manager. Closes scylladb/scylladb#20209	2024-09-12 16:42:28 +03:00
Nadav Har'El	d49dbb944c	Merge 'doc: move Alternator in the page tree and remove it's redundant ToC' from Anna Stuchlik This PR hides the ToC on the Alternator page, as we don't need it, especially at the end of the page. The ToC must be hidden rather than removed because removing it would, in turn, remove the "Getting Started With ScyllaDB Alternator" and "ScyllaDB Alternator for DynamoDB users" from the page tree and make them inaccessible. In addition, this PR moves Alternator higher in the page tree. Fixes https://github.com/scylladb/scylladb/issues/19823 Closes scylladb/scylladb#20565 * github.com:scylladb/scylladb: doc: move Alternator higher in the page tree doc: hide the redundant ToC on the Alternator page	2024-09-12 15:58:34 +03:00
Nadav Har'El	930accad12	alternator: return error on unused AttributeDefinitions A CreateTable request defines the KeySchema of the base table and each of its GSIs and LSIs. It also needs to give an AttributeDefinition for each attribute used in a KeySchema - which among other things specifies this attribute's type (e.g., S, N, etc.). Other, non-key, attributes do not have a specified type, and accordingly must not be mentioned in AttributeDefinitions. Before this patch, Alternator just ignored unused AttributeDefinitions entries, whereas DynamoDB throws an error in this case. This patch fixes Alternator's behavior to match DynamoDB's - and adds a test to verify this. Besides being more error-path-compatible with DynamoDB, this extra check can also help users: We already had one user complaining that an AttributeDefinitions setting he was using was ignored, not realizing that it wasn't used by any KeySchema. A clear error message would have saved this user hours of investigation. Fixes #19784. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20378	2024-09-12 15:37:18 +03:00
Pavel Emelyanov	632a65bffa	Merge 'repair: row_level: coroutinize more functions' from Avi Kivity Coroutinize more functions in row-level repair to improve maintainability. The functions all deal with repair buffers, so coroutinization does not affect performance. Cleanup, no reason to backport Closes scylladb/scylladb#20464 * github.com:scylladb/scylladb: repair: row_level: restore indentation repair: row_level: coroutinize repair_service::insert_repair_meta() repair: row_level: coroutinize repair_meta::get_full_row_hashes() repair: row_level: coroutinize repair_meta::apply_rows_on_follower() repair: row_level: coroutinize repair_meta::clear_working_row_buf() repair: row_level: coroutinize get_common_diff_detect_algorithm() repair: row_level: coroutinize repair_service::remove_repair_meta() (non-selective overload) repair: row_level: coroutinize repair_service::remove_repair_meta() (by-address overload) repair: row_level: coroutinize repair_service::remove_repair_meta() (by-id overload) repair: row_level: row_level_repair::run() repair: row_level: row_level_repair::send_missing_rows_to_follower_nodes() repair: row_level: row_level_repair::get_missing_rows_from_follower_nodes() repair: row_level: row_level_repair::negotiate_sync_boundary() repair: row_level: coroutinize repair_put_row_diff_with_rpc_stream_process_op() repair: row_level: coroutinize repair_meta::get_sync_boundary_handler() repair: row_level: coroutinize repair_meta::get_sync_boundary() repair: row_level: coroutinize repair_meta::repair_set_estimated_partitions_handler() repair: row_level: coroutinize repair_meta::repair_set_estimated_partitions() repair: row_level: coroutinize repair_meta::repair_get_estimated_partitions_handler() repair: row_level: coroutinize repair_meta::repair_get_estimated_partitions() repair: row_level: coroutinize repair_meta::repair_row_level_stop_handler() repair: row_level: coroutinize repair_meta::repair_row_level_stop() repair: row_level: coroutinize repair_meta::repair_row_level_start_handler() repair: row_level: coroutinize repair_meta::repair_row_level_start() repair: row_level: coroutinize repair_meta::get_combined_row_hash_handler() repair: row_level: coroutinize repair_meta::get_combined_row_hash() repair: row_level: coroutinize repair_meta::get_full_row_hashes_handler() repair: row_level: coroutinize repair_meta::get_full_row_hashes_with_rpc_stream() repair: row_level: coroutinize repair_meta::request_row_hashes()	2024-09-12 15:35:57 +03:00
Botond Dénes	f834ad81e0	docs/dev/reader-concurrency-semaphore.md: update the documentation on diagnostics dumps The part of the document which explains diagnostics dumps was due for an update. It was missing an explanation on the dumped stats and it also needs to explain the "Problematic permit" and "Identified bottleneck(s)".	2024-09-12 08:31:25 -04:00
Botond Dénes	fdff4beb1f	test/boost/reader_concurrency_semaphore_test: test the new diagnostics functionality Adjust the test reader_concurrency_semaphore_dump_reader_diganostics to also cover the new diagnostics functionality. The test is not a correctness test -- the output has to be inspected by a human. But it is good enough to make sure the code paths do not have any memory errors.	2024-09-12 08:31:25 -04:00
Botond Dénes	40b6616d3d	reader_concurrency_semaphore: add bottleneck self-diagnosis to diagnosis dump There are a few typical cases of bottlenecks, which can be easily identified when dumping the semaphore diagnostics. Identify and print these to fast-track investigations.	2024-09-12 08:31:25 -04:00
Botond Dénes	7d2b931619	reader_concurrency_semaphore: include trigger permit in diagnostic dump In the previous patch, we provided an opportunity for callers to provide a trigger permit, when calling `maybe_dump_reader_permit_diagnostics()`. If the caller provided the trigger permit, include its details in the dump, allowing the identification of the table and code-path of the permit which triggered the dump.	2024-09-12 08:30:50 -04:00
Kefu Chai	197451f8c9	utils/rjson.cc: include the function name in exception message recently, we are observing errors like: ``` stderr: error running operation: rjson::error (JSON SCYLLA_ASSERT failed on condition 'false', at: 0x60d6c8e 0x4d853fd 0x50d3ac8 0x518f5cd 0x51c4a4b 0x5fad446) ``` we only passed `false` to the `RAPIDJSON_ASSERT()` macro, so what we have is but the type of the error (rjson::error) and a backtrace. would be better if we can have more information without recompiling or fetching the debug symbols for decipher the backtrace. Refs scylladb/scylladb#20533 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20539	2024-09-12 15:22:49 +03:00
Anna Stuchlik	851e903f46	doc: move Alternator higher in the page tree	2024-09-12 14:08:26 +02:00
Anna Stuchlik	a32ff55c66	doc: hide the redundant ToC on the Alternator page This commit hides the ToC, as we don't need it, especially at the end of the page. The ToC must be hidden rather than removed because removing it would, in turn, remove the "Getting Started With ScyllaDB Alternator" and "ScyllaDB Alternator for DynamoDB users" from the page tree and make them inaccessible.	2024-09-12 14:01:15 +02:00
Benny Halevy	0b93409b44	cql_server: connection: process: fixup indentation Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-12 11:32:17 +02:00
Benny Halevy	71052dca6a	cql_server: connection: process_on_shard: drop permit parameter It is currently unused in `process_on_shard`, which generates an empty service_permit. The next patch may call process_on_shard in a loop, so it can't simply move the permit to the callee and better hold on to it until processing completes. `cql_server::connection::process` was turned into a coroutine in this patch to hold on to the permit parameter in a simple way. This is a preliminary step to changing `if (bounce_msg)` to `while (bounce_msg)` that will allow rebouncing the message in case it moved yet again when yielding in `process_on_shard`. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-12 11:32:17 +02:00
Benny Halevy	eb7fbdbed2	transport: server: pass bounce_to_shard as foreign shared ptr So it can safely passed between shards, as will be needed in the following patch that handles a (re)bounce_to_shard result from process_fn that's called by `process_on_shard` on the `move_to_shard`. With that in mind, pass the `bounce_to_shard` payload to `process_on_shard` rather than the foreign shared ptr since the latter grabs what it needs from it on entry and the shared_ptr can be released on the calling shard. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-12 11:32:15 +02:00
Benny Halevy	0df6f55379	cql_server: connection: process: add template concept for process_fn Quoting Avi Kivity: > Out of scope: we should consider detemplating this. As a follow-up we should consider that and pass a function object as process_fn, just make sure there are no drawbacks. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-12 10:26:13 +02:00
Benny Halevy	150dce5de0	cql_server: move process_fn_return_type to class definition So it can be used for a template concept in the next patch. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-12 10:26:13 +02:00
Alexey Novikov	8b6e987a99	test: add test_pinned_cl_segment_doesnt_resurrect_data add test for issue when writes in commitlog segments pinned to another table can be resurrected. This test based on dtest code published in #14870 and adapted for community version. It's a regression test for #15060 fix and should fail before this patch and succeed afterwards. Refs #14870, #15060 Closes scylladb/scylladb#20331	2024-09-12 10:58:22 +03:00
Takuya ASADA	90ab2a24df	toolchain: restore multiarch build When we introduced optimized clang at `6e487a4`, we dropped multiarch build on frozen toolchain, because building clang on QEMU emulation is too heavy. Actually, even after the patch merged, there are two mode which does not build clang, --clang-build-mode INSTALL_FROM and --clang-build-mode SKIP. So we should restore multiarch build only these mode, and keep skipping on INSTALL mode since it builds clang. Since we apply multiarch on INSTALL_FROM mode, --clang-archive replaced to --clang-archive-x86_64 and --clang-archive-aarch64. Note that this breaks compatibility of existing clang archive, since it changes clang root directory name from llvm-project to llvm-project-$ARCH. Closes #20442 Closes scylladb/scylladb#20444	2024-09-12 10:44:45 +03:00
Botond Dénes	c044904f07	reader_concurrency_semaphore: propagate permit to do_dump_reader_permit_diagnostics() Will be used in the next patch.	2024-09-12 00:51:56 -04:00
Botond Dénes	67565a5eee	reader_concurrency_semaphore: use consistent exception type for timeout When a read times out, we use different exception types for the permit's future (if the permit is waiting), or the permit's abort exception _ex (which is used to abort ongoing reads). This patch changes both to use named_semaphore_timed_out, which is the more verbose of the two.	2024-09-12 00:51:03 -04:00
Botond Dénes	036d27dc1b	reader_concurrency_semaphore: dump diagnostics when non-waiting reader times out Currently the semaphore only dumps diagnostics when a waiting reader times out. The diagnostics are also useful when a non-waiting reader (which is in the process of reading) times out, so also dump diagnostics in this case. Change the code to use a switch statement, so future addition of states don't miss updating this logic.	2024-09-12 00:51:03 -04:00
Kefu Chai	3e84d43f93	treewide: use seastar::format() or fmt::format() explicitly before this change, we rely on `using namespace seastar` to use `seastar::format()` without qualifying the `format()` with its namespace. this works fine until we changed the parameter type of format string `seastar::format()` from `const char*` to `fmt::format_string<...>`. this change practically invited `seastar::format()` to the club of `std::format()` and `fmt::format()`, where all members accept a templated parameter as its `fmt` parameter. and `seastar::format()` is not the best candidate anymore. despite that argument-dependent lookup (ADT for short) favors the function which is in the same namespace as its parameter, but `using namespace` makes `seastar::format()` more competitive, so both `std::format()` and `seastar::format()` are considered as the condidates. that is what is happening scylladb in quite a few caller sites of `format()`, hence ADT is not able to tell which function the winner in the name lookup: ``` /__w/scylladb/scylladb/mutation/mutation_fragment_stream_validator.cc:265:12: error: call to 'format' is ambiguous 265 \| return format("{} ({}.{} {})", _name_view, s.ks_name(), s.cf_name(), s.id()); \| ^~~~~~ /usr/bin/../lib/gcc/x86_64-redhat-linux/14/../../../../include/c++/14/format:4290:5: note: candidate function [with _Args = <const std::basic_string_view<char> &, const seastar::basic_sstring<char, unsigned int, 15> &, const seastar::basic_sstring<char, unsigned int, 15> &, const utils::tagged_uuid<table_id_tag> &>] 4290 \| format(format_string<_Args...> __fmt, _Args&&... __args) \| ^ /__w/scylladb/scylladb/seastar/include/seastar/core/print.hh:143:1: note: candidate function [with A = <const std::basic_string_view<char> &, const seastar::basic_sstring<char, unsigned int, 15> &, const seastar::basic_sstring<char, unsigned int, 15> &, const utils::tagged_uuid<table_id_tag> &>] 143 \| format(fmt::format_string<A...> fmt, A&&... a) { \| ^ ``` in this change, we change all `format()` to either `fmt::format()` or `seastar::format()` with following rules: - if the caller expects an `sstring` or `std::string_view`, change to `seastar::format()` - if the caller expects an `std::string`, change to `fmt::format()`. because, `sstring::operator std::basic_string` would incur a deep copy. we will need another change to enable scylladb to compile with the latest seastar. namely, to pass the format string as a templated parameter down to helper functions which format their parameters. to miminize the scope of this change, let's include that change when bumping up the seastar submodule. as that change will depend on the seastar change. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-11 23:21:40 +03:00
Pavel Emelyanov	f227f4332c	test: Remove unused path local variable Left after #20499 :( Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20540	2024-09-11 23:10:25 +03:00
Avi Kivity	ed7d352e7d	Merge 'Validate checksums for uncompressed SSTables' from Nikos Dragazis This PR introduces a new file data source implementation for uncompressed SSTables that will be validating the checksum of each chunk that is being read. Unlike for compressed SSTables, checksum validation for uncompressed SSTables will be active for scrub/validate reads but not for normal user reads to ensure we will not have any performance regression. It consists of: * A new file data source for uncompressed SSTables. * Integration of checksums into SSTable's shareable components. The validation code loads the component on demand and manages its lifecycle with shared pointers. * A new `integrity_check` flag to enable the new file data source for uncompressed SSTables. The flag is currently enabled only through the validation path, i.e., it does not affect normal user reads. * New scrub tests for both compressed and uncompressed SSTables, as well as improvements in the existing ones. * A change in JSON response of `scylla validate-checksums` to report if an uncompressed SSTable cannot be validated due to lack of checksums (no `CRC.db` in `TOC.txt`). Refs #19058. New feature, no backport is needed. Closes scylladb/scylladb#20207 * github.com:scylladb/scylladb: test: Add test to validate SSTables with no checksums tools: Fix typo in help message of scylla validate-checksums sstables: Allow validate_checksums() to report missing checksums test: Add test for concurrent scrub/validate operations test: Add scrub/validate tests for uncompressed SSTables test/lib: Add option to create uncompressed random schemas test: Add test for scrub/validate with file-level corruption test: Check validation errors in scrub tests sstables: Enable checksum validation for uncompressed SSTables sstables: Expose integrity option via crawling mutation readers sstables: Expose integrity option via data_consume_rows() sstables: Add option for integrity check in data streams sstables: Remove unused variable sstables: Add checksum in the SSTable components sstables: Introduce checksummed file data source implementation sstables: Replace assert with on_internal_error	2024-09-11 23:09:45 +03:00
Calle Wilund	b7839ec5d0	cql_test_env: Use temp socket + retry to ensure usable port for message_service if listen is enabled Fixes #20543 In cql_test_env, if cfg_in.ms_listen is set, we try to get a free port for the current test on which message service rpc can bind. This to allow multiple tests in parallel. However, we just do this by using random and getting a number, not actually verifying it against host ports in use. This is complicated further by the fact that port reuse is effectively disabled in seastar (see reactor::posix_reuseport_detect()). Due to this, the solution applied here is a combo of * Create temp socket with port = 0 to get a previously free port * Close socket right before listen (to handle reuse not working) * Retry on EADDRINUSE Closes scylladb/scylladb#20547	2024-09-11 23:02:41 +03:00
Aleksandra Martyniuk	31ea74b96e	db: system_keyspace: change version of topology_requests schema In `880058073b` a new column (request_type) was added to topology_requests table, but the table's schema version wasn't changed. Due to that during cluster upgrade, the old and the new versions occur but they are not distinguishable. Add offset to schema version of topology_requests table if it contains request_type column. Fixes: #20299. Closes scylladb/scylladb#20402	2024-09-11 16:36:35 +03:00
Piotr Dulikowski	d98708013c	Merge 'view: move view_build_status to group0' from Michael Litvak Migrate the `system_distributed.view_build_status` table to `system.view_build_status_v2`. The writes to the v2 table are done via raft group0 operations. The new parameter `view_builder_version` stored in `scylla_local` indicates whether nodes should use the old or the new table. New clusters use v2. Otherwise, the migration to v2 is initiated by the topology coordinator when the feature is enabled. It reads all the rows from the old table and writes them to the new table, and sets `view_builder_version` to v2. When the change is applied, all view_builder services are updated to write and read from the v2 table. The old table `system_distributed.view_build_status` is set to read virtually from the new table in order to maintain compatibility. When removing a node from the cluster, we remove its rows from the table atomically (fixes https://github.com/scylladb/scylladb/issues/11836). Also, during the migration, we remove all invalid rows. Fixes scylladb/scylladb#15329 dtest https://github.com/scylladb/scylla-dtest/pull/4827 Closes scylladb/scylladb#19745 * github.com:scylladb/scylladb: view: test view_build_status table with node replace test/pylib: use view_build_status_v2 table in wait_for_view view_builder: common write view_build_status function view_builder: improve migration to v2 with intermediate phase view: delete node rows from view_build_status on node removal view: sanitize view_build_status during migration view: make old view_build_status table a virtual table replica: move streaming_reader_lifecycle_policy to header file view_builder: test view_build_status_v2 storage_service: add view_build_status to raft snapshot view_builder: migration to v2 db:system_keyspace: add view_builder_version to scylla_local view_builder: read view status from v2 table view_builder: introduce writing status mutations via raft view_builder: pass group0_client and qp to view_builder view_builder: extract sys_dist status operations to functions db:system_keyspace: add view_build_status_v2 table	2024-09-11 13:02:58 +02:00
Nikos Dragazis	d1152a200f	test: Add test to validate SSTables with no checksums In a previous patch we extended the return status of `sstables::validate_checksums()` to report if an SSTable cannot be validated due to a missing CRC component (i.e., CRC.db does not appear in TOC.txt). Add a test case for this. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:40 +03:00
Nikos Dragazis	1f275c71b1	tools: Fix typo in help message of scylla validate-checksums Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:39 +03:00
Nikos Dragazis	5c0a7f706b	sstables: Allow validate_checksums() to report missing checksums Change the return type of `sstable::validate_checksums()` from binary (valid/invalid) to a ternary (valid/invalid/no_checksums). The third status represents uncompressed SSTables without a CRC component (no entry for CRC.db in the TOC). Also, change the JSON response of `sstable validate-checksums` to expose the new status. Replace the boolean value for valid/invalid checksums with an object that contains two boolean keys: one that indicates if the SSTable has checksums, and one that indicates if the checksums are valid or not. The second key is optional and appears only if the SSTable has checksums. Finally, update the documentation to reflect the changes in the API. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:39 +03:00
Nikos Dragazis	5a284f4a9d	test: Add test for concurrent scrub/validate operations Theoretically it is possible to launch more than one scrub instances simultaneously. Since the checksum component is a shared resource, accesses have to be synchronized. Add a test that launches two scrub operations in validate mode and ensures that the checksum component is loaded once, referenced by all scrub instances via shared pointers, and deleted once the scrub operations finish. Introduce an injection point to achieve concurrent execution of scrubs. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:39 +03:00
Nikos Dragazis	e2353f3b3e	test: Add scrub/validate tests for uncompressed SSTables Currently the unit tests check scrub in validate mode against compressed SSTables only. Mirror the tests for uncompressed SSTables as well. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:39 +03:00
Nikos Dragazis	2991b09c8e	test/lib: Add option to create uncompressed random schemas Extend the `random_schema_specification` to support creating both compressed and uncompressed schemas. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:32 +03:00
Nikos Dragazis	4f56c587f6	test: Add test for scrub/validate with file-level corruption Currently, we test scrub/validate only against a corrupted SSTable with content-level corruption (out-of-order partition key). Add a test for file-level corruption as well. This should trigger the checksum check in the underlying compressed file data source implementation. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:28:59 +03:00
Nikos Dragazis	cc10a5f287	test: Check validation errors in scrub tests Scrub was extended in PR #11074 to report validation errors but the unit tests were not updated. Update the tests to check the validation errors reported by scrub. Validation errors must be zero for valid SSTables and non-zero for invalid SSTables. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:28:59 +03:00
Nikos Dragazis	719757fba9	sstables: Enable checksum validation for uncompressed SSTables Extend the `sstable::validate()` to validate the checksums of uncompressed SSTables. Given that this is already supported for compressed SSTables, this allows us to provide consistent behavior across any type of SSTable, be it either compressed or uncompressed. The most prominent use case for this is scrub/validate, which is now able to detect file-level corruption in uncompressed SSTables as well. Note that this change will not affect normal user reads which skip checksum validation altogether. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:28:59 +03:00
Nikos Dragazis	716fc487fd	sstables: Expose integrity option via crawling mutation readers Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:28:59 +03:00
Nikos Dragazis	1d2dc9f2e1	sstables: Expose integrity option via data_consume_rows() Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:28:59 +03:00
Nikos Dragazis	2feced32f7	sstables: Add option for integrity check in data streams Add a new boolean parameter in `sstable::data_stream()` to enable/disable integrity mechanisms in the underlying data streams. Currently, this only affects uncompressed SSTables and it allows to enable/disable checksum validation on each chunk. The validation happens transparently via the checksummed data source implementation. The reason we need this option is to allow differentiating the behavior between normal user reads and scrub/validate reads. We would like to enable scrub to verify checksums for uncompressed SSTables, while leaving normal user reads unchanged for performance reasons (read amplification due to round up of reads to chunk size and loading of the CRC component). Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:27:54 +03:00
Nikos Dragazis	d5bd40ad2c	sstables: Remove unused variable Remove unused stream variable from `sstable::data_stream()`. This was introduced in commit `47e07b787e` but never used. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:27:54 +03:00
Nikos Dragazis	2575d20f41	sstables: Add checksum in the SSTable components Uncompressed SSTables store their checksums in a separate CRC.db file. Add this in the list of SSTable components. Since this component is used only for validation, load the component on-demand for validation tasks and delete it when all validation tasks finish. In more detail: - Make the checksum component shareable and weakly referencable. Also, add a constructor since it is no longer an aggregate. - Use a weak pointer to store a non-owning reference in the components and a shared pointer to keep the object alive while validation runs. Once validation finishes, the component should be cleaned up automatically. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:27:38 +03:00
Nikos Dragazis	b7dfba4c18	sstables: Introduce checksummed file data source implementation Introduce a new data source implementation for uncompressed SSTables. This is just a thin wrapper for a raw data source that also performs checksum validation for each chunk. This way we can have consistent behavior for compressed and uncompressed SSTables. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 12:26:18 +03:00
Botond Dénes	0e5b444777	Merge 'database::get_all_tables_flushed_at: fix return value' from Lakshmi Narayanan Sreethar The `database::get_all_tables_flushed_at` method returns a variable without setting the computed all_tables_flushed_at value. This causes its caller, `maybe_flush_all_tables` to flush all the tables everytime regardless of when they were last flushed. Fix this by returning the computed value from `database::get_all_tables_flushed_at`. Fixes #20301 Requires a backport to 6.0 and 6.1 as they have the same issue. Closes scylladb/scylladb#20471 * github.com:scylladb/scylladb: cql-pytest: add test to verify compaction_flush_all_tables_before_major_seconds config database::get_all_tables_flushed_at: fix return value	2024-09-11 11:43:45 +03:00
Amnon Heiman	46792bd04f	docs/alternator/compatibility.md: explain the consumed capacity provisioned This patch change the alternator documentation to express that the provisoned units are stored and return but Alternator ignores them. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-09-10 17:28:31 -04:00
Amnon Heiman	3726c20564	Add test/alternator/test_provisioned_throughput.py The test_provisioned_throughput.py test ProvisionedThroughput support. The first test, check that ProvisionedThroughput can be set and get when using describe table. The second test check that missing read or write will throw an exception. The third test check that when using billing PAY_PER_REQUEST it returns zero for the read and write units. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-09-10 17:27:19 -04:00
Amnon Heiman	9b5f29b6bc	test/alternator/util.py: Allow override BillingMode This patch adds the ability to override the BillingMode. If a BillingMode is provided to the create_test_table function, it will override the default BillingMode. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-09-10 17:06:47 -04:00
Amnon Heiman	c76347032d	alternator/executor.cc: Store ProvisionedThroughput This patch adds the ability to store and retrieve the ProvisionedThroughput in a table. The information is stored in the table tags. We use the TTL convention used in alternator, and the tags will be: system:provisioned_rcu and system:provisioned_wcu. verify_billing_mode function now return a struct with the billing mode information. The code of describe_table now check if the provision tags exists and return the RCU and WCU accordingly. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-09-10 17:06:40 -04:00
Benny Halevy	4e8f3f4cdd	cql-pytest: add test_compaction_tombstone_gc Test tombstone garbage collection with: 1. conflicting live data in memtable (verifying there is no regression in this area) 2. deletion in memtable (reproducing scylladb/scylladb#20423) 3. materialized view update in memtable (reproducing scylladb/scylladb#20424) in materialized_views Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:06:23 +03:00
Benny Halevy	9270348c38	sstable_compaction_test: add mv_tombstone_purge_test Simulate view updates pattern and verify that they don't inhibit tombstone garbage collection. Verify fix for scylladb/scylladb#20424 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:06:23 +03:00
Benny Halevy	0407e50aa4	sstable_compaction_test: tombstone_purge_test: test that old deleted data do not inhibit tombstone garbage collection Tests fix for scylladb/scylladb#20423 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:06:06 +03:00
Benny Halevy	a7caa79df7	sstable_compaction_test: tombstone_purge_test: add testlog debugging Add some testlog debug printouts for the make_* helpers. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:58 +03:00
Benny Halevy	470d301fe3	sstable_compaction_test: tombstone_purge_test: make_expiring: use next_timestamp Rather than forging a timestamp from the gc_clock just use `next_timestamp` do it can be considered for tomebstone purging purposes. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:58 +03:00
Benny Halevy	5849ba83e0	sstable, compaction: add debug logging for extended min timestamp stats Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Benny Halevy	7d893a5ed9	compaction: get_max_purgeable_timestamp: use memtable and sstable extended timestamp stats When purging regular tombstone consult the min_live_timestamp, if available. For shadowable_tombstones, consult the min_memtable_live_row_marker_timestamp, if available, otherwise fallback to the min_live_timestamp. If both are missing, fallback to the legacy (and inaccurate) min_timestamp. Fixes scylladb/scylladb#20423 Fixes scylladb/scylladb#20424 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Benny Halevy	57e9e9c369	compaction: define max_purgeable_fn Before we add a new, is_shadowable, parameter to it. And define global `can_always_purge` and `can_never_purge` functions, a-la `always_gc` and `never_gc`. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Benny Halevy	b6fabd98c6	tombstone: can_gc_fn: move declaration to compaction_garbage_collector.hh And define `never_gc` globally, same as `always_gc` Before adding a new, is_shadowable parameter to it. Since it is used in the context of compaction it better fits compaction_garbage_collector header rather than tombstone.hh Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Benny Halevy	4de4af954f	sstables: scylla_metadata: add ext_timestamp_stats Store and retrieve the optional extended timestamp statistics (min_live_timestamp and min_live_row_marker_timestamp) in the scylla_metadata component. Note that there is no need for a cluster feature to store those attributes since the scylla_metadata on-disk format is extensible so that old sstables can be read by new versions, seeing the extra stats is missing, and new sstables can be read by old versions that ignore unknown scylla metadata section types. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Benny Halevy	6f202cf48b	compaction_group, storage_group, table_state: add extended timestamp stats getters To return the minimum live timestamp and live row-marker timestamp across a compaction_group, storage_group, or table_state. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Benny Halevy	14d86a3a12	sstables, memtable: track live timestamps When garbage collecting tombstones, we care only about shadowing of live data. However, currently we track min/max timestamp of both live and dead data, but there is no problem with purging tombstones that shadow dead data (expired or shdowed by other tombstones in the sstable/memtable). Also, for shadowable tombstones, we track live row marker timestamps separately since, if the live row marker timestamp is greater than a shadowable tombstone timestamp, then the row marker would shadow the shadowable tombstone thus exposing the cells in that row, even if their timestasmp may be smaller than the shadow tombstone's. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:49 +03:00
Abhi	9b09439065	raft: Add descriptions for requested abort errors Fixes: scylladb/scylladb#18902 Closes scylladb/scylladb#20291	2024-09-10 17:56:29 +02:00
Botond Dénes	de81388edb	Merge 'commitlog: Handle oversized entries' from Calle Wilund Refs #18161 Yet another approach to dealing with large commitlog submissions. We handle oversize single mutation by adding yet another entry typo: fragmented. In this case we only add a fragment (aha) of the data that needs storing into each entry, along with metadata to correlate and reconstruct the full entry on replay. Because these fragmented entries are spread over N segments, we also need to add references from the first segment in a chain to the subsequent ones. These are released once we clear the relevant cf_id count in the base. * This approach has the downside that due to how serialization etc works w.r.t. mutations, we need to create an intermediate buffer to hold the full serialized target entry. This is then incrementally written into entries of < max_mutation_size, successively requesting more segments. On replay, when encountering a fragment chain, the fragment is added to a "state", i.e. a mapping of currently processing frag chains. Once we've found all fragments and concatenated the buffers into a single fragmented one, we can issue a replay callback as usual. Note that a replay caller will need to create and provide such a state object. Old signature replay function remains for tests and such. This approach bumps the file format (docs to come). To ensure "atomicity" we both force synchronization, and should the whole op fail, we restore segment state (rewinding), thus discarding data all we wrote. Closes scylladb/scylladb#19472 * github.com:scylladb/scylladb: commitlog/database: Make some commitlog options updatable + add feature listener features/config: Add feature for fragmented commitlog entries docs: Add entry on commitlog file format v4 commitlog_test: Add more oversized cases commitlog_replayer: Replay segments in order created commitlog_replayer: Use replay state to support fragmented entries commitlog_replayer: coroutinize partly commitlog: Handle oversized entries	2024-09-10 17:15:46 +03:00
Benny Halevy	8d67357c42	memtable_encoding_stats_collector: update row_marker: do nothing if missing If the row_marker is missing then its timestamp is missing as well, so there's no point calling update_timestamp for it. Better return early. This should cause no functional change. The following patch will add more logic for tracking extended timestamp stats. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 16:46:34 +03:00
Pavel Emelyanov	b6f662417c	table: Remove unused database& argument from take_snapshot() method Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20496	2024-09-10 14:53:06 +03:00
Gleb Natapov	af83c5e53e	group0: stop group0 before draining storage service during shutdown Currently storage service is drained while group0 is still active. The draining stops commitlogs, so after this point no more writes are possible, but if group0 is still active it may try to apply commands which will try to do writes and they will fail causing group0 state machine errors. This is benign since we are shutting down anyway, but better to fix shutdown order to keep logs clean. Fixes scylladb/scylladb#19665	2024-09-10 13:15:56 +02:00
Lakshmi Narayanan Sreethar	a0f4fe3fc4	cql-pytest: add test to verify compaction_flush_all_tables_before_major_seconds config Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-10 16:39:05 +05:30
Lakshmi Narayanan Sreethar	4ca720f0bd	database::get_all_tables_flushed_at: fix return value The `database::get_all_tables_flushed_at` method returns a variable without setting the computed all_tables_flushed_at value. This causes its caller, `maybe_flush_all_tables` to flush all the tables everytime regardless of when they were last flushed. Fix this by returning the computed value from `database::get_all_tables_flushed_at`. Fixes #20301 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-10 16:35:47 +05:30
Yaniv Michael Kaul	a4ff0aae47	HACKIGN.md: clarify the use of dbuild when running test.py If you are using dbuild, that's where test.py needs to run. Also, replace 'Docker image' with the more generic 'container' term. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Closes scylladb/scylladb#20336	2024-09-10 13:40:45 +03:00
Botond Dénes	08f109724b	docs/cql/ddl.rst: fix description of sstable_compression ScyllaDB doesn't support custom compressors. The available compressors are the only available ones, not the default ones. Adjust the text to reflect this. Closes scylladb/scylladb#20225	2024-09-10 13:39:24 +03:00
Pavel Emelyanov	cfa59ab73d	test: Use single temp dir for sharded<sstables::test_env> The test-env in question is mostly started in one-shard mode. Also there are several boost tests that start sharded<> environment. In that case instances on different shards live in different temp dirs. That's not critical yet, but better to have single directory for the whole test. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20412	2024-09-10 11:25:04 +03:00
Artsiom Mishuta	f95c257a1e	[test.py]: Fail test teardown in case of task leakage In test.py every asyncio task spawned during the test must be finished before the next test, otherwise, tests might affect each other results. The developers are responsible for writing asyncio code in a way that doesn’t leave task objects unfinished. Test.py has a mechanism that helps test writers avoid such tasks. At the end of each test case, it verifies that the test did not produce/leave any tasks and sets an event object that fails the next test at the start if this is the case(issue https://github.com/scylladb/scylladb/issues/16472) The problem with this was that breaking the next test was counterintuitive, and the logging for this situation was insufficient and unobvious. notes: Task.cancel() is not an option to avoid task leakage 1) Calling cancel() Does Not Cancel The Task : the cancel() method just request that the target task cancel. 2) Calling cancel() Does Not Block Until The Task is Cancelled: If the caller needs to know the task is cancelled and done, it could await for the target 3) In particular PR, task.cancel() cancell task on client(ManagerClient) but not on http server(ScyllaManager). so "await" is needed. Closes scylladb/scylladb#20012	2024-09-10 10:51:45 +03:00
Pavel Emelyanov	ac2127a640	test: Call table::make_sstable() directly in compaction test The test in question generates a bunch of table_for_tests objects and creates sstables for each. For that it calls test_env::make_sstable(), but it can be made shorter, by calling table method directly. The hidden goal of this change is to remove the explicit caller of table::dir() method. The latter is going away. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20451	2024-09-10 10:19:20 +03:00
Botond Dénes	76bb22664a	Merge 'Sanitize open_sstables() helper in compaction test' from Pavel Emelyanov This includes - coroutinization - elimination of unused overload Closes scylladb/scylladb#20456 * github.com:scylladb/scylladb: test: Squash two open_sstables() helper together test: Coroutinize open_sstables() helper	2024-09-10 10:18:33 +03:00
Botond Dénes	a4a4797e27	Merge 'Alternator: tests and other preparation towards allowing adding a GSI to an existing table' from Nadav Har'El This series prepares us for working on #11567 - allow adding a GSI to a pre-existing table. This will require changing the implementation of GSIs in Alternator to not use real columns in the schema for the materialized view, and instead of a computed column - a function which extracts the desired member from the `:attrs` map and de-serializes it. This series does not contain the GSI re-implementation itself. Rather it contains a few small cleanups and mostly - new regression tests that cover this area, of adding and removing a GSI, and using a GSI, in more details than the tests we already had. I developed most of these tests while working on buggy fixes for #11567; The bugs in those implementations were exposed by the tests added here - they exposed bugs both in the new feature of adding or removing a GSI, and also regressions to the ordinary operation of GSI. So these tests should be helpful for whoever ends up fixing #11567, be it me based on my buggy implementation (which is _not_ included in this patch series), or someone else. No backports needed - this is part of a new feature, which we don't usually backport. Closes scylladb/scylladb#20383 * github.com:scylladb/scylladb: test/alternator: more extensive tests for GSI with two new key attributes test/alternator: test invalid key types for GSI test/alternator: test combination of LSI and GSI test/alternator: expand another test to use different write operations test/alternator: test GSIs with different key types alternator: better error message in some cases of key type mismatch test/alternator: test for more elaborate GSI updates test/alternator: strengthen tests for empty attribute values test/alternator: fix typo in test_batch.py test/alternator: more checks for GSI-key attribute validation Alternator: drop unneeded "IS NOT NULL" clauses in MV of GSI/LSI test/alternator: add more checks for adding/deleting a GSI test/alternator: ensure table deletions in test_gsi.py	2024-09-10 10:13:52 +03:00
Pavel Emelyanov	42f8d06a17	test: Use correct schema in directory tests with created table There are some test cases in sstable_directory_test test actually create a table with CQL and then try to manipulate its sstables with the help of sstable_directory. Those tests use existing local helper that starts sharded<sstable_directory> and this helper passes test-local static schema to sstable_directory constructor. As a result -- the schema of a table that test case created and the schema that sstable_directory works with are different. They match in the columns layout, which helps the test cases pass, but otherwise are two different schema objects with different IDs. It's more correct to use table schema for those runs. The fix introduces another helper to start sharded<sstable_directory>, and the older wrapper around cql_test_env becomes unused. Drop it too not to encourage future tests use it and re-introduce schema mismatch again. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20499	2024-09-10 09:56:26 +03:00
Benny Halevy	f47b5e60bc	sstable_directory: create_pending_deletion_log: place pending_delete log under the base directory To be able to atomically delete sstables both in base table directory and in its sub-directories, like `staging/`, use a shared pending_delete_dir under under the base directory. Note that this requires loading and processing the base directory first. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 09:28:13 +03:00
Benny Halevy	44bd183187	sstables: storage: keep base directory in base class so we can use the base (table) directory for e.g. pending_delete logs, in the next patch. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 09:28:13 +03:00
Benny Halevy	027e64876a	sstables: storage: define opened_directory in header file So it can be used outside the storage module in the following patches. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 09:28:13 +03:00
Benny Halevy	a7b92d7b6f	sstable_directory: use only dirlog Currently, there are leftover log messages using sstlog rather than dirlog, that was introduced in `aebd965f0e`, and that makes debugging harder. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 09:28:11 +03:00
Botond Dénes	fc690a60d8	Update tools/cqlsh submodule * tools/cqlsh 86a280a1...b09bc793 (6): > build(deps): bump actions/download-artifact in /.github/workflows > cqlshlib/test: Add test_formatting.py > cqlshlib/test: Use assertEqual instead of assertEquals > cqlsh.py: Send DESCRIBE statement to server before parsing > cqlsh.py: Fix indentation > cqlsh.py: change shebang to /usr/bin/env python3	2024-09-10 08:11:40 +03:00
Lakshmi Narayanan Sreethar	2148e33d37	compaction: remove unnecessary share bump for split, scrub, and upgrade When split, scrub, and upgrade compactions ran under the compaction group, they had to bump up their shares to a minimum of 200 to prevent slow progress as they neared completion, especially in workloads with inconsistent ingestion rates. Since commit `e86965c2` moved these compactions to the maintenance group, this share bump is no longer necessary. This patch removes the unnecessary share allocation. Fixes #20224 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com> Closes scylladb/scylladb#20495	2024-09-09 22:03:38 +03:00
Avi Kivity	9448260b30	Merge 'major compaction: check only sstables being compacted for tombstone garbage collection' from Lakshmi Narayanan Sreethar Any expired tombstone can be garbage collected if it doesn't shadow data in the commit log, memtable, or uncompacting SSTables. This PR introduces a new mode to major compaction, enabled by the `consider_only_existing_data` flag that bypasses these checks. When enabled, memtables and old commitlog segments are cleared with a system-wide flush and all the sstables (after flush) are included in the compaction, so that it works with all data generated up to a given time point. This new mode works with the assumption that newly written data will not be shadowed by expired tombstones. So it ignores new sstables (and new data written to memtable) created after compaction started. Since there was a system wide flush, commitlog checks can also be skipped when garbage collecting tombstones. Introducing data shadowed by a tombstone during compaction can lead to undefined behavior, even without this PR, as the tombstone may or may not have already been garbage collected. Fixes #19728 Closes scylladb/scylladb#20031 * github.com:scylladb/scylladb: cql-pytest: add test to verify consider_only_existing_data compaction option tools/scylla-nodetool: add consider-only-existing-data option to compact command api: compaction: add `consider_only_existing_data` option compaction: consider gc_check_only_compacting_sstables when deducing max purgeable timestamp compaction: do not check commitlog if gc_check_only_compacting_sstables is enabled tombstone_gc_state: introduce with_commitlog_check_disabled() compaction: introduce new option to check only compacting sstables for gc compaction: rename maybe_flush_all_tables to maybe_flush_commitlog compaction: maybe_flush_all_tables: add new force_flush param	2024-09-09 20:45:41 +03:00
Avi Kivity	894b85ce95	Merge 'hints: send hints with CL=ALL if target is leaving' from Piotr Dulikowski Currently, when attempting to send a hint, we might choose its recipients in one of two ways: - If the original destination is a natural endpoint of the hint, we only send the hint to that node and none other, - Otherwise, we send the hint to all current replicas of the mutation. There is a problem when we decommission a node: while data is streamed away from that node, it is still considered to be a natural endpoint of the data that it used to own. Because of that, it might happen that a hint is sent directly to it but streaming will miss it, effectively resulting in the hint being discarded. As sending the hint _only_ to the leaving replica is a rather bad idea, send the hint to all replicas also in the case when the original destination of the hint is leaving. Note that this is a conservative fix written only with the decommission + vnode-based keyspaces combo in mind. In general, such "data loss" can occur in other situations where the replica set is changing and we go through a streaming phase, i.e. other topology operations in case of vnodes and tablet load balancing. However, the consistency guarantees of hinted handoff in the face of topology changes are not defined and it is not clear what they should be, if there should be any at all. The picture is further complicated by the fact that hints are used by materialized views, and sending view updates to more replicas than necessary can introduce inconsistencies in the form of "ghost rows". This fix was developed in response to a failing test which checked the hint replay + decommission scenario, and it makes it work again. Fixes scylladb/scylla-dtest#4582 Refs scylladb/scylladb#19835 Should be backported to 6.0 and 6.1; the dtest started failing due to topology on raft, which sped up execution of the test and exposed the preexisting problem. Closes scylladb/scylladb#20488 * github.com:scylladb/scylladb: test: topology_custom/test_hints: consistency test for decommission test: topology_custom/test_hints: move sync point helpers to top level test: topology/util: extract find_server_by_host_id hints: send hints with CL=ALL if target is leaving hints: inline do_send_one_mutation	2024-09-09 18:23:13 +03:00
Avi Kivity	c3e19425bd	Merge 'docs/dev/docker-hub.md: refresh aio-max-nr calculation' from Laszlo Ersek ~~~ What we have today in "docs/dev/docker-hub.md" on "aio-max-nr" dates back to scylla commit `f4412029f4` ("docs/docker-hub.md: add quickstart section with --smp 1", 2020-09-22). Problems with the current language: - The "65K" claim as default value on non-production systems is wrong; "fs/aio.c" in Linux initializes "aio_max_nr" to 0x10000, which is 64K. - The section in question uses equal signs (=) incorrectly. The intent was probably to say "which means the same as", but that's not what equality means. - In the same section, the relational operator "<" is bogus. The available AIO count must be at least as high (>=) as the requested AIO count. - Clearer names should be used; adjust_max_networking_aio_io_control_blocks() in "src/core/reactor.cc" sets a great example: - "reactor::max_aio" should be called "storage_iocbs", - "detect_aio_poll" should be called "preempt_iocbs", - "reactor_backend_aio::max_polls" should be called "network_iocbs". - The specific value 10000 for the last one ("network_iocbs") is not correct in scylla's context. It is correct as the Seastar default, but scylla has used 50000 since commit `2cfc517874` ("main, test: adjust number of networking iocbs", 2021-07-18). Rewrite the section to address these problems. See also: - https://github.com/scylladb/scylladb/issues/5981 - https://github.com/scylladb/seastar/pull/2396 - https://github.com/scylladb/scylladb/pull/19921 Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> ~~~ No need for backporting; the documentation being refreshed targets developers as audience, not end-users. Closes scylladb/scylladb#20398 * github.com:scylladb/scylladb: docs/dev/docker-hub.md: refresh aio-max-nr calculation docs/dev/docker-hub.md: strip trailing whitespace	2024-09-09 15:04:38 +03:00
Botond Dénes	3e0bff161c	Merge 'Use yielding directory lister in sstable_directory' from Pavel Emelyanov The yielding lister is considered to be better replacement that scan_dir(lambda) one. Also, the sstable directory will be patched to scan the contents of S3 bucket and yielding lister fits better for generalization. Closes scylladb/scylladb#20114 * github.com:scylladb/scylladb: sstable_directory: Fix indentation after previous patches sstable_directory: Use yielding lister in .handle_sstables_pending_delete() sstable_directory: Use yielding lister in .cleanup_column_family_temp_sst_dirs() sstable_directory: Use yielding lister in .prepare() sstable_directory: Shorten lister loop sstable_directory: Use with_closeable() in .process() directory_lister: Add noexcept default move-constructor	2024-09-09 14:35:51 +03:00
Pavel Emelyanov	0f48847d02	test: Use shorter with_sstable_directory overload() In sstable directory test there are two of those -- one that works on path, state, env and callback, and the other one that just needs env and callback, getting path from env and assuming state is normal. Two test cases in this test can enjoy the shorter one. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20395	2024-09-09 14:25:24 +03:00
Pavel Emelyanov	2bfbbaffac	test: Use sstables::test_env to make sstables for schema loader test This test calls manager directly, but it's shorter to ask test_env for that Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20431	2024-09-09 14:22:58 +03:00
Takuya ASADA	e36c939505	dist: tune LimitNOFILES for large nodes On very large node, LimitNOFILES=80000 may not enough size, it can cause "Too many files" error. To avoid that, let's increase LimitNOFILES on scylla_setup stage, generate optimal value calurated from memory size and number of cpus. Closes scylladb/scylla-enterprise#4304 Closes scylladb/scylladb#20443	2024-09-09 14:13:49 +03:00
Piotr Smaron	60af48f5fd	cql: fix exception when validating KS in CREATE TABLE `c70f321c6f` added an extra check if KS exists. This check can throw `data_dictionary::no_such_keyspace` exception, which is supposed to be caught and a more user-friendly exception should be thrown instead. This commit fixes the above problem and adds a testcase to validate it doesn't appear ever again. Also, I moved the check for the keyspace outside of the `for` loop, as it doesn't need to be checked repeatedly. Fixes: scylladb/scylladb#20097 Closes scylladb/scylladb#20404	2024-09-09 13:30:57 +03:00
Nadav Har'El	ee7d4d8825	test/alternator: more extensive tests for GSI with two new key attributes The case of a GSI with two key attributes (hash and range) which were both not keys in the base table is a special case, not supported by CQL but allowed in Alternator. We have several tests for this case, but they don't cover all the strange possibilities that a GSI row disappears / reappears when one or two of the attributes is updated / inserted / deleted. So this patch includes a more extensive test for this case. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:49 +03:00
Nadav Har'El	ad53d6a230	test/alternator: test invalid key types for GSI This patch adds a test that types which are not allowed for GSI keys - basically any type except S(tring), B(ytes) or N(number), are rejected as expected - an error path that we didn't cover in existing tests. The new test passes - Alternator doesn't have a bug in this area, and as usual, also passes on DynamoDB. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:49 +03:00
Nadav Har'El	c4021d0819	test/alternator: test combination of LSI and GSI To allow adding a GSI to an existing table (refs #11567), we plan to re-implement GSIs to stop forcing their key attribute to become a real column in the schema - and let it remains a member of the map ":attrs" like all non-key attributes. But since LSIs can only be defined on table creation time, we don't have to change the LSI implementation, and these can still force their key to become a real column. What the test in this patch does is to verify that using the same attribute as a key of both GSI and LSI on the same table works. There's a high risk that it won't work: After all, the LSI should force the attribute to become a real column (to which base reads and writes go), but the GSI will use a computed column which reads from ":attrs", no? Well, it turns out that view.cc's value_getter::operator() always had a surprising exception which "rescues" this test and makes it pass: Before using a computed column, this code checks if a base-table column with the same name exists, and if it does, it is used instead of the computed column! It's not clear why this logic was chosen, but it turns out to be really useful for making the test in this test pass. And it's important that if we ever change that unintuitive behavior, we will have this test as a regression test. The new test unsurprisingly passes on current Scylla because its implementation of GSI and LSI is still the same. But it's an important regression test for when we change the GSI implementation. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:49 +03:00
Nadav Har'El	7563d0a8a1	test/alternator: expand another test to use different write operations Expand another Alternator test (test_gsi.py::test_gsi_missing_attribute) to write items not just using PutItem, but also using UpdateItem and BatchWriteItem. There is a risk that these different operations use slightly different code paths - so better check all of them and not just PutItem. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:49 +03:00
Nadav Har'El	4d02beec53	test/alternator: test GSIs with different key types All of the tests in test/alternator/test_gsi.py use strings as the GSI's keys. This tests a lot of GSI functionality, but we implicitly assumed that our implementation used an already-correct and already-tested implementation of key columns and MV, which if it works for one type, works for other types as well. This assumption will no longer hold if we reimplement GSI on a "computed column" implementation, which might run different code for different types of GSI key attributes (the supported types are "S"tring, "B"ytes, and "N"umber). So in this patch we add tests for writing and reading different types of GSI key attributes. These tests showed their importance as regression tests when the first draft of the GSI reimplementation series failed them. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:49 +03:00
Nadav Har'El	80a0798e77	alternator: better error message in some cases of key type mismatch Alternator uses a common function get_typed_value() to read the values of key attribute and confirm they have the expected type (key attributes have a fixed type in the schema). If the type is wrong, we want to print a "Type mismatch" error message. But the current implementation did the checks in the wrong order, and as a result could print a "Malformed value object" message instead of a "Type mismatch". That could happen if the wrong type is a boolean, map, list, or basically any type whose JSON representation is not a string. The allowed key types - bytes), string and number - all have string representations in JSON, but still we should first report the mismatched type and only report the "Malformed object" if the type matches but the JSON is faulty. In addition to fixing the error message, we fix an existing test which complained in a comment (but ignored) that the error message in some case (when trying to use a map where a key is expected) the strange "Malformed value object" instead of the expected "Type mismatch". The next patch will add an additional reproducer for this problem and its fix. That test will do: ``` with pytest.raises(ClientError, match='ValidationException.*mismatch'): test_table_gsi_6.put_item(Item={'p': p, 's': True}) ``` I.e., it tries to set a boolean value for a string key column, and expect to get the "Type mismatch" error and not the ugly "Malformed value object". Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:49 +03:00
Nadav Har'El	624ed32278	test/alternator: test for more elaborate GSI updates Most tests in test_gsi.py involve simple updates to a GSI, just creating a GSI row. Although a couple of tests did involve more complex operations (such as an update requiring deleting an old row from the GSI and inserting a new one,), we did not have a single organized test designed to check all these cases, so we add one in this patch. This test (test_update_gsi_pk) will be important for verifying the low-level implementation of the new GSI implementation that we plan to based on computed columns. Early versions of that code passed many of the simpler tests, but not this one. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:49 +03:00
Nadav Har'El	65d4ddf093	test/alternator: strengthen tests for empty attribute values We soon plan to refactor Alternator's GSI and change the validation of values set in attributes which are GSI keys. It's important to test that when updating attributes that are not GSI keys - and are either base- table keys or normal non-key attributes - the validation didn't change. For example, empty strings are still not allowed in base-table key attributes, but are allowed (since May 2020 in DynamoDB) in non-key attributes. We did have tests in this area, but this patch strengthens them - adding a test for non-key attribute, and expanding the key-attribute test to cover the UpdateItem and BatchWriteItem operations, not just PutItem. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 13:14:41 +03:00
Avi Kivity	9a5061209f	Merge '[test.py] Enable allure for python test' from Andrei Chekun To enhance the test reports UX: 1. switching off/on passed/failed/skipped test for better visibility 2. better searching in test results 3. understanding the trends of execution for each test 4. better configurability of the final report Enable allure adapter for all python tests. Add tags and parameters to the test to be able to distinguish them across modes and runs. Related: https://github.com/scylladb/qa-tasks/issues/1665 Related: https://github.com/scylladb/scylladb/pull/19335 Related: https://github.com/scylladb/scylladb/pull/18169 Closes scylladb/scylladb#19942 * github.com:scylladb/scylladb: [test.py] Clean duplicated arg for test suite [test.py] Enable allure for python test	2024-09-09 12:53:00 +03:00
Nadav Har'El	5859daed68	test/alternator: fix typo in test_batch.py Two tests had a typo 'item' instead of 'Item'. If Scylla had a bug, this could have caused these tests to miss the bug. Scylla passes also the fixed test, because Scylla's behavior is correct. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 12:09:25 +03:00
Nadav Har'El	1f8e39f680	test/alternator: more checks for GSI-key attribute validation When an attribute is a GSI key, DynamoDB imposes certain rules when writing values for it - it must be of the declared type for that key, and can't be an empty string. We had tests for this, but all of them did the write using the PutItem operation. In this patch we also test the same things using the UpdateItem and BatchWriteItem operations. Because Scylla has different code paths for these three operations, and each code path needs to remember to call the validation function, all three should all be checked and not just PutItem. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 12:09:25 +03:00
Nadav Har'El	cf5d7ce212	Alternator: drop unneeded "IS NOT NULL" clauses in MV of GSI/LSI Scylla's materialized views naturally skip any base rows where the view's key isn't set (is NULL), because we can't create a view row with a null key. To make the user aware that this is happening, the user is required to add "WHERE ... IS NOT NULL" for the view's key columns when defining the view. However, the only place that these extra IS NOT NULL clauses are checked are in the CQL "CREATE MATERIALIZED VIEWS" statement - they are completely ignored in all other places in the code. In particular, when we create a materialized view in Alternator (GSI or LSI), we don't have to add these "IS NOT NULL" clauses, as they are outright ignored. We didn't know they were ignored, and made an effort to add them - but no matter how incorrectly we did it, it didn't matter :-) In commit `2bf2ffd3ed` it turned out we had a typo that caused the wrong column name to be printed. Also, even today we are still missing base key columns that aren't listed as a view key in Alternator but still added as view clustering keys in Scylla - and again the fact these were missing also didn't matter. So I think it's time to stop pretending, and stop calculating these "IS NOT NULL" strings, so this patch outright removes them from the Alternator view-creation code. Beyond being a nice cleanup of unnecessary and inaccurate code, it will also be necessary when we allow in later patches to index for an Alternator attribute "x" not a real column x in the base table but rather an element in the ":attrs" map - so adding a "x IS NOT NULL" isn't only unnecessary, it is outright illegal: The expression evaluation code, even though it doesn't do anything with the "IS NOT NULL" expression, still verifies that "x" is a valid column, which it isn't. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 12:09:25 +03:00
Nadav Har'El	8beaa9d10e	test/alternator: add more checks for adding/deleting a GSI We already have tests for the feature of adding or removing a GSI from an existing table, which Alternator doesn't yet support (issue #11567). In this patch we add another check, how after a GSI is added, you can no longer add items with the wrong type for the indexed type, and after removing a GSI, you can. The expanded tests pass on DynamoDB, and obviously still xfail on Alternator because the feature is not yet implemented. Refs #11567. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 12:09:25 +03:00
Nadav Har'El	ce19311ab3	test/alternator: ensure table deletions in test_gsi.py Most of the Alternator tests are careful to unconditionally remove the test tables, even if the test fails. This is important when testing on a shared database (e.g., DynamoDB) but also useful to make clean shutdown faster as there should be no user table to flush. We missed a few such cases in test_gsi.py, and fixed some of them in commit `59c1498338` but still missed a few, and this patch fixes some more instances of this problem. We do this by using the context manager new_test_table() - which automatically deletes the table when done - instead of the function create_test_table() which needs an explicit delete at the end. There are no functional changes in this patch - most of the lines changed are just reindents. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-09-09 12:09:25 +03:00
Kefu Chai	ccbd3eb9f7	main: do not register redis and alternator services if not enabled in main.cc, we start redis with `ss.local().register_protocol_server()` only if it is enabled. but `storage_service` always calls `stop_server()` with _all_ registered server, no matter if they have started or not. in general, it does not hurt. for instance, `redis::controller::stop_server()` is a noop, if the controller is not started. but `storage_service` still print the logging message like: ``` INFO 2024-09-04 11:20:02,224 [shard 0:main] storage_service - Shutting down redis server INFO 2024-09-04 11:20:02,224 [shard 0:main] storage_service - Shutting down redis server was successful ``` this could be confusing or at least distracting when a field engineer looks at the log. also, please note, `redis_port` and `redis_ssl_port` cannot be changed dynamically once scylla server is up, so we do not need to worry about "what if the redis server is started at runtime, how can is be stopped?". the same applies to alternator service. in this change, to avoid surprises, we conditionally register the protocol servers with the storage service based on their enabled statuses. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20472	2024-09-09 08:44:50 +03:00
Avi Kivity	58713f3080	types: remove some unused free functions These functions are unused, so safe to remove, and reduce the work to convert to managed_bytes{,_view}. Closes scylladb/scylladb#20482	2024-09-09 08:36:33 +03:00
Kefu Chai	720997d1de	cql3/statements: mark format string as `constexpr const` after switching over to the new `seastar::format()` which enables the compile-time format check, the fmt string should be a constexpr, otherwise `fmt::format()` is not able to perform the check at compile time. to prepare for bumping up the seastar module to a version which contains the change of `seastar::format()`, let's mark the format string with `constexpr const`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20484	2024-09-09 08:35:45 +03:00
Piotr Dulikowski	6f3d0af994	test: topology_custom/test_hints: consistency test for decommission Adds the test_hints_consistency_during_decommission test which reproduces the failure observed in scylladb/scylla-dtest#4582. It uses error injections, including the newly added topology_coordinator_pause_after_streaming injection, to reliably orchestrate the scenario observed there. In a nutshell, the test makes sure to replay hints after streaming during decommission has finished, but before the cluster switches to reading from new replicas. Without the fix, hints would be replayed to the decommissioned node and then would be lost forever after the cluster start reading from new replicas.	2024-09-08 10:51:38 +02:00
Piotr Dulikowski	30d53167c9	test: topology_custom/test_hints: move sync point helpers to top level Move create_sync_point and await_sync_point from the scope of the test_sync_point test to the file scope. They will be used in a test that will be introduced in the commit that follows.	2024-09-08 10:51:38 +02:00
Piotr Dulikowski	a75d0c0bfa	test: topology/util: extract find_server_by_host_id Move it out from test_mv_tablets_replace.py. It will be used by a test introduced in a later commit.	2024-09-08 10:51:38 +02:00
Piotr Dulikowski	61ac0a336d	hints: send hints with CL=ALL if target is leaving Currently, when attempting to send a hint, we might choose its recipients in one of two ways: - If the original destination is a natural endpoint of the hint, we only send the hint to that node and none other, - Otherwise, we send the hint to all current replicas of the mutation. There is a problem when we decommission a node: while data is streamed away from that node, it is still considered to be a natural endpoint of the data that it used to own. Because of that, it might happen that a hint is sent directly to it but streaming will miss it, effectively resulting in the hint being discarded. As sending the hint _only_ to the leaving replica is a rather bad idea, send the hint to all replicas also in the case when the original destiantion of the hint is leaving. Note that this is a conservative fix written only with the decommission + vnode-based keyspaces combo in mind. In general, such "data loss" can occur in other situations where the replica set is changing and we go through a streaming phase, i.e. other topology operations in case of vnodes and tablet load balancing. However, the consistency guarantees of hinted handoff in the face of topology changes are not defined and it is not clear what they should be, if there should be any at all. The picture is further complicated by the fact that hints are used by materialized views, and sending view updates to more replicas than necessary can introduce inconsistencies in the form of "ghost rows". This fix was developed in response to a failing test which checked the hint replay + decommission scenario, and it makes it work again. Fixes scylladb/scylla-dtest#4582 Refs scylladb/scylladb#19835	2024-09-08 10:50:59 +02:00
Piotr Dulikowski	8abb06ab82	hints: inline do_send_one_mutation It's a small method and it is only used once in send_one_mutation. Inlining it lets us get rid of its declaration in the header - now, if one needs to change the variables passed from one function to another, it is no longer necessary to change the header.	2024-09-08 07:19:35 +02:00
Avi Kivity	ab32ce6b45	Merge 'Coroutinize sstable::read_summary() method' from Pavel Emelyanov Shorter and simpler this way. Hopefully it doesn't sit on critical paths Closes scylladb/scylladb#20460 * github.com:scylladb/scylladb: sstables: Fix indentation after previous patch sstables: Coroutinize sstable::read_summary()	2024-09-06 18:45:54 +03:00
Pavel Emelyanov	103c68b419	sstables: Restore indentation after previous patch Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-06 18:24:56 +03:00
Pavel Emelyanov	c47c0f1cd6	sstables: Coroutinize remove_unshared_sstables() Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-06 18:24:40 +03:00
Kefu Chai	aeaeaf345d	compaction: use structured binding when appropriate for better readability Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20473	2024-09-06 18:17:48 +03:00
Kamil Braun	427ad2040f	Merge 'test: randomized failure injection for Raft-based topology' from Evgeniy Naydanov The idea of the test is to have a cluster where one node is stressed with injections and failures and the rest of the cluster is used to make progress of the raft state machine. To achieve this following two lists introduced in the PR: - ERROR_INJECTIONS in error_injections.py - CLUSTER_EVENTS in cluster_events.py Each cluster event is an async generator which has 2 yields and should be used in the following way: 0. Start the generator: ```python >>> cluster_event_steps = cluster_event(manager, random_tables, error_injection) ``` 1. Run the prepare part (before the first yield) ```python >>> await anext(cluster_event_steps) ``` 2. Run the cluster event itself (between the yields) ```python >>> await anext(cluster_event_steps) ``` 3. Run the check part (after the second yield) ```python >>> await anext(cluster_event, None) ``` Closes scylladb/scylladb#16223 * github.com:scylladb/scylladb: test: randomized failure injection for Raft-based topology test: error injections for Raft-based topology [test.py] topology.util: add get_non_coordinator_host() function [test.py] random_tables: add UDT methods [test.py] random_tables: add CDC methods [test.py] api: get scylla process status [test.py] api: add expected_server_up_state argument to server_add()	2024-09-06 14:00:41 +02:00
Pavel Emelyanov	226fd03bae	Merge 'service/qos: remove unused marked_for_deletion field from service_level struct' from Piotr Dulikowski The `service_level::marked_for_deletion` field is always set to `false`. It might have served some purpose in the past, but now it can be just removed, simplifying the code and eliminating confusion about the field. This is just code cleanup, no backport is needed. Closes scylladb/scylladb#20452 * github.com:scylladb/scylladb: service/qos: remove the marked_for_deletion parameter service/qos: add constructors to service_level	2024-09-06 11:44:25 +03:00
Kamil Braun	52fdf5b4c9	test: test_raft_no_quorum: increase raft timeout in debug mode The test cases in this file use an error injection to reduce raft group 0 timeouts (from the default 1 minute), in order to speed up the tests; the scenarios expect these timeouts to happen, so we want them to happen as quick as possible, but we don't want to reduce timeouts so much that it will make other operations fail when we don't expect them to (e.g. when the test wants to add a node to the cluster). Unfortunately the selected 5 seconds in debug mode was not enough and made the tests flaky: scylladb/scylladb#20111. Increase it to 10 seconds. This unfortunately will slow down these tests as they have to sometimes wait for 10 seconds for the timeout to happen. But better to have this than a flaky test. Fixes: scylladb/scylladb#20111 Closes scylladb/scylladb#20320	2024-09-06 11:40:09 +03:00
Avi Kivity	384a09585b	repair: row_level: repair_get_row_diff_with_rpc_stream_process_op: simplify return value During review of `0857b63259` it was noticed that the function repair_get_row_diff_with_rpc_stream_process_op() and its _slow_path callee only ever return stop_iteration::no (or throw an exception). As such, its return value is useless, and in fact the only caller ignores it. Simplify by returning a plain future<>. Closes scylladb/scylladb#20441	2024-09-06 11:39:21 +03:00
Kefu Chai	034c1df29b	auth/authentication_options: move fmt::formatter up so that it is accessible from its caller. if we enforce the compile-time format string check, the formatter would need the access to the specialization of `fmt::formatter` of the arguments being foramtted. to be prepared for this change, let's move the `fmt::formatter` specialization up, otherwise we'd have following error after switching to the compile-time format string check introduced by a recent seastar change: ``` In file included from ./auth/authenticator.hh:22: ./auth/authentication_options.hh:50:49: error: call to consteval function 'fmt::basic_format_string<char, auth::authentication_option &>::basic_format_string< char[32], 0>' is not a constant expression 50 \| : std::invalid_argument(fmt::format("The {} option is not supported.", k)) { \| ^ ./auth/authentication_options.hh:57:13: error: explicit specialization of 'fmt::formatter<auth::authentication_option>' after instantiation 57 \| struct fmt::formatter<auth::authentication_option> : fmt::formatter<string_view> { \| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ /usr/include/fmt/base.h:1228:17: note: implicit instantiation first required here 1228 \| -> decltype(typename Context::template formatter_type<T>().format( \| ^ In file included from replica/distributed_loader.cc:30: ``` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20447	2024-09-06 09:12:38 +03:00
Pavel Emelyanov	527fc9594a	sstables: Fix indentation after previous patch And move the comment inside if while at it, it looks better in there (and makes less churn in the patch itself) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-06 08:43:08 +03:00
Pavel Emelyanov	f7325586f3	sstables: Coroutinize sstable::read_summary() Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-06 08:43:07 +03:00
Evgeniy Naydanov	dd99cf197d	test: randomized failure injection for Raft-based topology The idea of the test is to have a small cluster, where one node is stressed with injections and failures and the rest of the cluster is used to make progress of the Raft state machine. To achieve this following two lists introduced in the commit: - ERROR_INJECTIONS in error_injections.py - CLUSTER_EVENTS in cluster_events.py Each cluster event is an async generator which has 2 yields and should be used in the following way: 0. Start the generator: >>> cluster_event_steps = cluster_event(manager, random_tables, error_injection) 1. Run the prepare part (before the first yield) >>> await anext(cluster_event_steps) 2. Run the cluster event itself (between the yields) >>> await anext(cluster_event_steps) 3. Run the check part (after the second yield) >>> await anext(cluster_event, None)	2024-09-05 22:11:32 +00:00
Evgeniy Naydanov	769424723b	test: error injections for Raft-based topology Add following error injections: - stop_after_init_of_system_ks - stop_after_init_of_schema_commitlog - stop_after_starting_gossiper - stop_after_starting_raft_address_map - stop_after_starting_migration_manager - stop_after_starting_commitlog - stop_after_starting_repair - stop_after_starting_cdc_generation_service - stop_after_starting_group0_service - stop_after_starting_auth_service - stop_during_gossip_shadow_round - stop_after_saving_tokens - stop_after_starting_gossiping - stop_after_sending_join_node_request - stop_after_setting_mode_to_normal_raft_topology - stop_before_becoming_raft_voter - topology_coordinator_pause_after_updating_cdc_generation - stop_before_streaming - stop_after_streaming - stop_after_bootstrapping_initial_raft_configuration	2024-09-05 22:11:31 +00:00
Evgeniy Naydanov	ac4ffbad5c	[test.py] topology.util: add get_non_coordinator_host() function Add get_non_coordinator_host() function which returns ServerInfo for the first host which is not a coordinator or None if there is no such host. Also rework get_coordinator_host() to not fail if some of the hosts don't have a host id.	2024-09-05 22:11:31 +00:00
Evgeniy Naydanov	d95d698601	[test.py] random_tables: add UDT methods Add .add_udt() / .drop_udt() methods.	2024-09-05 22:11:31 +00:00
Evgeniy Naydanov	8cb442ca50	[test.py] random_tables: add CDC methods Add .enabled_cdc() / .disable_cdc() methods.	2024-09-05 22:11:31 +00:00
Evgeniy Naydanov	a7119cf420	[test.py] api: get scylla process status Add `server_get_process_status(server_id)` API call and wait_for_scylla_process_status() helper function.	2024-09-05 22:11:31 +00:00
Evgeniy Naydanov	241bbb4172	[test.py] api: add expected_server_up_state argument to server_add() Allow to return from server_add() when a server reaches specified state. One of: - PROCESS_STARTED - HOST_ID_QUERIED (previously called NOT_CONNECTED) - CQL_CONNECTED (renamed from CONNECTED) - CQL_QUERIED (was just QUERIED) Also, rename CqlUpState to ServerUpState and move to internal_types.	2024-09-05 22:11:31 +00:00
Pavel Emelyanov	f02a686115	schema: Ditch make_shared_schema() helper Now it's unused Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 19:34:00 +03:00
Pavel Emelyanov	d045aa6df7	test: Tune up indentation in uncompressed_schema() After it was switched to use schema builder, the indenation of untouched lines deserves one extra space. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 19:33:29 +03:00
Pavel Emelyanov	a1deba0779	test: Make tests use schema_builder instead of make_shared_schema Everything, but perf test is straightforward switch. The perf-test generated regular columns dynamically via vector, with builder the vector goes away. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 19:31:30 +03:00
Avi Kivity	c57b8dd0bf	repair: row_level: restore indentation	2024-09-05 18:38:43 +03:00
Avi Kivity	710977ef88	repair: row_level: coroutinize repair_service::insert_repair_meta() Some of the indentation was broken, and is partially repaired by this change.	2024-09-05 17:59:42 +03:00
Avi Kivity	f23a32ed84	repair: row_level: coroutinize repair_meta::get_full_row_hashes()	2024-09-05 17:56:27 +03:00
Avi Kivity	607747beb1	repair: row_level: coroutinize repair_meta::apply_rows_on_follower()	2024-09-05 17:55:07 +03:00
Avi Kivity	89d4394d12	repair: row_level: coroutinize repair_meta::clear_working_row_buf()	2024-09-05 17:52:32 +03:00
Pavel Emelyanov	69a5ec69c4	test: Use table storage options in sstable_directory_test When creating sstables this test allocates temporary local options. That works, because this test doesn't run on object storage, but it's more correct to pick storage options from the table at hand. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20440	2024-09-05 17:48:25 +03:00
Avi Kivity	4cfc25f8d7	repair: row_level: coroutinize get_common_diff_detect_algorithm() The function is threaded, but the inner lambda can be coroutinized.	2024-09-05 17:47:27 +03:00
Michael Litvak	9545e0a114	view: test view_build_status table with node replace Add a test replacing a node and verifying the contents of the view_build_status table are updated as expected, having rows for the new node and no rows for the old node.	2024-09-05 15:42:35 +03:00
Michael Litvak	3ca5dd537f	test/pylib: use view_build_status_v2 table in wait_for_view Change the util function wait_for_view to read the view build status from the system.view_build_status_v2 table which replaces system_distributed.view_build_status. The old table can still be used but it is less efficient because it's implemented as a virtual table which reads from the v2 table, so it's better to read directly from the v2 table. This can cause slowness in tests. The additional util function wait_for_view_v1 reads from the old table. This may be needed in upgrade tests if the v2 table is not available yet.	2024-09-05 15:42:35 +03:00
Michael Litvak	5c95aaae0d	view_builder: common write view_build_status function When writing to the view_build_status we have common logic related to upgrade and deciding whether to write to sys_dist ks or group0. Move this common logic to a generic function used by all functions writing to the table.	2024-09-05 15:42:35 +03:00
Michael Litvak	c1f3517a75	view_builder: improve migration to v2 with intermediate phase Add an intermediate phase to the view builder migration to v2 where we write to both the old and new table in order to not lose writes during the migration. We add an additional view builder version v1_5 between v1 and v2 where we write to both tables. We perform a barrier before moving to v2 to ensure all the operations to the old table are completed.	2024-09-05 15:42:35 +03:00
Michael Litvak	446ad3c184	view: delete node rows from view_build_status on node removal When a node is removed we want to clean its rows from the view_build_status table. Now when removing a node and generating the topology state update, we generate also the mutations to delete all the possible rows belonging to the node from the table.	2024-09-05 15:42:35 +03:00
Michael Litvak	08462aaff7	view: sanitize view_build_status during migration When migrating the view_build_status to v2, skip adding any leftover rows that don't correspond to an existing node or an existing view. Previously such rows could have been created and not cleaned, for example when a node is removed.	2024-09-05 15:42:35 +03:00
Michael Litvak	78d6ff6598	view: make old view_build_status table a virtual table After migrating the view build status from system_distributed.view_build_status to system.view_build_status_v2, we set system_distributed.view_build_status to be a virtual table, such that reading from it is actually reading from the underlying new table. The reason for this is that we want to keep compatibility with the old table, since it exists also in Cassandra and it is used by various external tools to check the view build status. Making the table virtual makes the transition transparent for external users. The two tables are in different keyspaces and have different shard mapping. The v1 table is a distributed table with a normal shard mapping, and the v2 table is a local table using the null sharder. The virtual reader works by constructing a multishard reader which reads the rows from shard zero, and then filtering it to get only the rows owned by the current shard.	2024-09-05 15:42:35 +03:00
Michael Litvak	09eadcff08	replica: move streaming_reader_lifecycle_policy to header file move the class streaming_reader_lifecycle_policy to a header file in order to make it reusable in other places.	2024-09-05 15:42:35 +03:00
Michael Litvak	22f4f1fa49	view_builder: test view_build_status_v2 Add tests to verify the new view_build_status_v2 is used by the view_builder and can be read from all nodes with the expected values. Also test a migration from the v1 layout to v2.	2024-09-05 15:42:35 +03:00
Michael Litvak	fcf66ad541	storage_service: add view_build_status to raft snapshot Include the table system.view_build_status_v2 in the raft snapshot, and also the view_builder version parameter.	2024-09-05 15:42:30 +03:00
Michael Litvak	8d25a4d678	view_builder: migration to v2 Migrate view_builder to v2, to store the view build status of all nodes in the group0 based table view_build_status_v2. Introduce a feature view_build_status_on_group0 so we know when all nodes are ready to migrate and use the new table. A new cluster is initialized to use v2. Otherwise, The topology coordinator initiates the migration when the feature is enabled, if it was not done already. The migration reads all the rows in the v1 table and writes it via group0 to the v2 table, together with a mutation that updates the view_builder parameter in scylla_local to v2. When this mutation is applied, it updates the view_builder service to start using the v2 table.	2024-09-05 15:41:04 +03:00
Michael Litvak	f3887cd80b	db:system_keyspace: add view_builder_version to scylla_local Add a new scylla_local parameter view_builder_version, and functions to read and mutate the value. The version value defaults to v1 if it doesn't exist in the table.	2024-09-05 15:41:04 +03:00
Michael Litvak	d58a8930c4	view_builder: read view status from v2 table Update the view_status function to read from the new view_build_status_v2 table when enabled. The code to read and extract the values is identical to v1 and v2 except it accesses different keyspace and table, so the common code is extracted to the view_status_common function and used by both v1 and v2 flows with appropriate parameters.	2024-09-05 15:41:04 +03:00
Michael Litvak	05d18b818f	view_builder: introduce writing status mutations via raft Introduce the announce_with_raft function as alternative to writing view build status mutations to the table in system_distributed. Instead, we can apply the mutations via group0 operation to the view_build_status_v2 table. All the view_builder functions that write to the view_build_status table can be configured by a flag to either write the legacy way or via raft.	2024-09-05 15:41:04 +03:00
Michael Litvak	b8c7a10ae6	view_builder: pass group0_client and qp to view_builder Store references of group0_client and query_processor in the view_builder service. They are required for generating mutations and writing them via group0.	2024-09-05 15:41:04 +03:00
Michael Litvak	b2332c5a72	view_builder: extract sys_dist status operations to functions Extract all the update and read operations of a view build status in the table system_distributed.view_build_status to separate functions.	2024-09-05 15:41:04 +03:00
Michael Litvak	bf4a58bf91	db:system_keyspace: add view_build_status_v2 table add the table system.view_build_status_v2 with the same schema as system_distributed.view_build_status.	2024-09-05 15:41:04 +03:00
Gleb Natapov	807e37502a	db/consistency_level: do not use result from heat weighted load balancer if it contains duplicates Because of https://github.com/scylladb/scylladb/issues/9285 heat weighted load balancer may sometimes return same node twice. It may cause wrong data to be read or unexpected errors to be returned to a client. Since the original bug is not easy to fix and it is rare lets introduce a workaround. We will check for duplicates and will use non HWLB one if one is found. Fixes scylladb/scylladb#20430 Closes scylladb/scylladb#20414	2024-09-05 15:21:35 +03:00
Wojciech Mitros	c1b0434c16	test: finish mv view update explicitly instead of relying on delay duration When testing mv admission control, we perform a large view update and check if the following view update can be admitted due to the high view backlog usage. We rely on a delay which keeps the backlog high for longer to make sure the backlog is still increased during the second write. However, in some test runs the delay is not long enough, causing the second write to miss the large backlog and not hit admission control. In this patch we keep the increased backlog high using another injection instead of relying on a delay to make absolute sure that the backlog is still high during the second write. Fixes scylladb/scylladb#20382 Closes scylladb/scylladb#20445	2024-09-05 15:08:04 +03:00
Lakshmi Narayanan Sreethar	7c5efab7d5	cql-pytest: add test to verify consider_only_existing_data compaction option Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:34:13 +05:30
Lakshmi Narayanan Sreethar	68a902f74a	tools/scylla-nodetool: add consider-only-existing-data option to compact command Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:34:06 +05:30
Lakshmi Narayanan Sreethar	84d06a13c7	api: compaction: add `consider_only_existing_data` option Added a new parameter `consider_only_existing_data` to major compaction API endpoints. When enabled, major compaction will: - Force-flush all tables. - Force a new active segment in the commit log. - Compact all existing SSTables and garbage-collect tombstones by only checking the SSTables being compacted. Memtables, commit logs, and other SSTables not part of the compaction will not be checked, as they will only contain newer data that arrived after the compaction started. The `consider_only_existing_data` is passed down to the compaction descriptor's `gc_check_only_compacting_sstables` option to ensure that only the existing data is considered for garbage collection. The option is also passed to the `maybe_flush_commitlog` method to make sure all the tables are flushed and a new active segment is created in the commit log. Fixes #19728 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:25:45 +05:30
Lakshmi Narayanan Sreethar	98bc44f900	compaction: consider gc_check_only_compacting_sstables when deducing max purgeable timestamp When gc_check_only_compacting_sstables is enabled, get_max_purgeable_timestamp should not check memtables and other sstables that are not part of the compaction to deduce the max purgeable timestamp. Refs #19728 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:25:45 +05:30
Lakshmi Narayanan Sreethar	7b9ce8e040	compaction: do not check commitlog if gc_check_only_compacting_sstables is enabled When the compaction_descriptor's gc_check_only_compacting_sstables flag is enabled, create and pass a copy of the get_tombstone_gc_state that will skip checking the commitlog. Refs #19728 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:25:45 +05:30
Lakshmi Narayanan Sreethar	12fa40154b	tombstone_gc_state: introduce with_commitlog_check_disabled() Added a new method, `with_commitlog_check_disabled`, that returns a new copy of the tombstone_gc_state but with commitlog check disabled. This will be used by a following patch to disable commitlog checks during compaction. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:25:45 +05:30
Lakshmi Narayanan Sreethar	5b8c6a8a5e	compaction: introduce new option to check only compacting sstables for gc Added new option, `gc_check_only_compacting_sstables`, to compaction_descriptor to control the garbage collection behavior. The subsequent patches will use this flag to decide if the garbage collection has to check only the SSTables being compacted to collect tombstones. This option is disabled for now and will be enabled based on a new compaction parameter that will be added later in this patch series. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:25:45 +05:30
Lakshmi Narayanan Sreethar	5e6bffc146	compaction: rename maybe_flush_all_tables to maybe_flush_commitlog Major compaction flushes all tables as a part of flushing the commitlog. After forcing new active segments in the commitlog, all the tables are flushed to enable reclaim of older commitlog segments. The main goal is to flush the commitlog and flushing all the table is just a dependency. Rename maybe_flush_all_tables to maybe_flush_commitlog so that it reflects the actual intent of the major compaction code. Added a new wrapper method to database::flush_all_tables(), database::flush_commitlog(), that is now called from maybe_flush_commitlog. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:25:45 +05:30
Lakshmi Narayanan Sreethar	fa2488cc83	compaction: maybe_flush_all_tables: add new force_flush param Add a new parameter, `force_flush` to the maybe_flush_all_tables() method. Setting `force_flush` to true will flush all the tables regardless of when they were flushed last. This will be used by the new compaction option in a following patch. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-09-05 17:25:45 +05:30
Laszlo Ersek	53524974db	docs/dev/maintainer.md: clarify "Updating submodule references" Before the introduction of "scripts/refresh-submodules.sh", there was indeed some manual work for the maintainer to do, hence "publish your work" must have sounded correct. Today, the phrase "publish your work" sounds confusing. Commit `71da4e6e79` ("docs: Document sync-submodules.sh script in maintainer.md", 2020-06-18) should have arguably reworded the last step of the submodule refresh procedure; let's do it now. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#20333	2024-09-05 13:57:32 +03:00
Pavel Emelyanov	1f0db29ef6	test: Remove unused directory semaphore The with_sstable_dir() helper no longer needs one, it used to pass it as argument to sstable_directory constructor, but now the directory doesn't need it (takes semaphore via table object). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20396	2024-09-05 13:11:35 +03:00
Kefu Chai	b4fc24cc1f	github: use needs.read-toolchain.outputs.image for build-scylla so we don't need to hardwire the image on which we build scylla. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20370	2024-09-05 12:58:36 +03:00
Pavel Emelyanov	955391d209	sstable_directory: Fix indentation after previous patches Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 11:19:19 +03:00
Pavel Emelyanov	2febde24f3	sstable_directory: Use yielding lister in .handle_sstables_pending_delete() Indentation is deliberately left broken. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 11:19:19 +03:00
Pavel Emelyanov	02aac3e407	sstable_directory: Use yielding lister in .cleanup_column_family_temp_sst_dirs() Indentation is deliberately left broken Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 11:19:19 +03:00
Pavel Emelyanov	ff77a677a6	sstable_directory: Use yielding lister in .prepare() Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 11:19:19 +03:00
Pavel Emelyanov	7b5fe6bee6	sstable_directory: Shorten lister loop Squash call to lister.get() and check for the returned value into while()'s condition. This saves few more lines of code as well. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 11:19:19 +03:00
Pavel Emelyanov	5dc266cefa	sstable_directory: Use with_closeable() in .process() The method already uses yielding lister, but handles the exceptions explicitly. Use with_closeable() helper, it makes the code shorter. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 11:19:19 +03:00
Pavel Emelyanov	7742b90cb1	directory_lister: Add noexcept default move-constructor It's required to make it possible to push lister into with_closeable(). Its requiremenent of nothrow-move-constructible doesn't accept default-generated one. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 11:10:21 +03:00
Nikos Dragazis	2450afb934	sstables: Replace assert with on_internal_error The `skip()` method of the compressed data source implementation uses an assert statement to check if the given offset is valid. Replace this with `on_internal_error()` to fail gracefully. An invalid offset shouldn't bring the whole server down. Also, enhance the error message for unsynced compressed readers. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-05 11:03:54 +03:00
Pavel Emelyanov	da598a6210	test: Restore indentation after previous changes Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:38:01 +03:00
Pavel Emelyanov	e16c07c896	test: Threadify tombstone_in_tombstone2() Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	28d016f312	test: Threadify range_tombstone_reading() Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	7d567d07ad	test: Threadify tombstone_in_tombstone() Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	a34e38f070	test: Threadify broken_ranges_collection() Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	eac4ec47f8	test: Threadify compact_storage_dense_read() Indentation is deliberately left broken. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	322c1ee9c5	test: Threadify compact_storage_simple_dense_read() Indentation is deliberately left broken. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	df71b3e446	test: Threadify compact_storage_sparse_read() Indentation is deliberately left broken. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	142ccc64fb	test: Simplify test_range_reads() counting It used to keep counter with the help of a smart pointer, now it can just use on-stack variable. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	a78ab2e998	test: Simplify test_range_reads() inner loop It used to rely on bool (wrapped with pointer) and future<>-based loop helper, now it can just break from the while loop. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	c84ae64562	test: Threadify test_range_reads() itself And update its callers again. Preserve no longer relevant local smart pointers until next patch. Indentation is deliberately left broken. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:33 +03:00
Pavel Emelyanov	253d53b6a1	test: Threadify test_range_reads() callers Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:36:00 +03:00
Pavel Emelyanov	fd8bb0c46c	test: Threadify generate_clustered() itself And update its callers again. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:35:59 +03:00
Pavel Emelyanov	f500ee690b	test: Threadify generate_clustered() callers Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:34:54 +03:00
Pavel Emelyanov	08186c048d	test: Threadify test_no_clustered test And update its callers. Indentation is deliberately left broken. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:26:25 +03:00
Pavel Emelyanov	5f0a40f959	test: Threadify nonexistent_key test Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 10:26:13 +03:00
Pavel Emelyanov	a150a63259	test: Squash two open_sstables() helper together One accepts integer generations, another one accepts "generic" ones. The latter is only called by the former, so no sense in keeping it around. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 09:08:40 +03:00
Pavel Emelyanov	4184c688ea	test: Coroutinize open_sstables() helper Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-05 09:08:12 +03:00
Piotr Dulikowski	ecd53db3b0	service/qos: remove the marked_for_deletion parameter It is always set to false and it doesn't seem to serve any function now.	2024-09-04 21:52:34 +02:00
Piotr Dulikowski	bae6076541	service/qos: add constructors to service_level Add a default constructor and a constructor which explicitly initializes all fields of the service_level structure. This is done in order to make sure that removal of the marked_for_deletion field can be done safely - otherwise, for example, service_level could be aggregate-initialized with an incomplete list of values for the fields, and removing marked_for_deletion which is in the middle of the struct would cause the is_static field to be initialized with the value that was designated for marked_for_deletion. As a bonus, make sure that marked_for_deletion and is_static bool fields are initialized in the default constructor to false in order to avoid potential undefined behavior.	2024-09-04 21:52:13 +02:00
Avi Kivity	ec8590ae6c	Merge 'Always pass `abort_source&` to `raft_group0_client::hold_read_apply_mutex`' from Kamil Braun There are two versions of `raft_group0_client::hold_read_apply_mutex`, one takes `abort_source&`, the other doesn't. Modify all call sites that used the non-abort-source version to pass an `abort_source&`, allowing us to remove the other overload. If there is no explicit reason not to pass an `abort_source&`, then one should be passed by default -- it often prevents hangs during shutdown. --- No backport needed -- no known issues affected by this change. Closes scylladb/scylladb#19996 * github.com:scylladb/scylladb: raft_group0_client: remove `hold_read_apply_mutex` overload without `abort_source&` storage_service: pass `_abort_source` to `hold_read_apply_mutex` group0_state_machine: pass `_abort_source` to `hold_read_apply_mutex` api: move `reload_raft_topology_state` implementation inside `storage_service`	2024-09-04 21:35:27 +03:00
Kefu Chai	fe0e961856	docs: do not install scylla/ppa repo when perform upgrade for following reasons: 1. the ppa in question does not provide the build for the latest ubuntu's LTS release. it only builds for trusty, xenial, bionic and jammy. according to https://wiki.ubuntu.com/Releases, the latest LTS release is ubuntu noble at the time of writing. 2. the ppa in question does not provide the packages used in production. it does provides the package for building scylla 3. after we introduced the relocatable package, there is no need to provide extra user space dependencies apart from scylla packages. so, in this change, we remove all references to enabling the Scylla/PPA repository. Fixes scylladb/scylladb#20449 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20450	2024-09-04 20:30:40 +03:00
Avi Kivity	20b79816f1	repair: row_level: coroutinize repair_service::remove_repair_meta() (non-selective overload)	2024-09-04 18:43:19 +03:00
Avi Kivity	3b9ac51b6b	repair: row_level: coroutinize repair_service::remove_repair_meta() (by-address overload)	2024-09-04 18:39:21 +03:00
Avi Kivity	704e3f5432	repair: row_level: coroutinize repair_service::remove_repair_meta() (by-id overload)	2024-09-04 18:37:48 +03:00
Avi Kivity	9612c4d790	repair: row_level: row_level_repair::run() The function itself is threaded, but the inner lambdas are coroutinized (except one which is expected to run in a thread, and so is threaded).	2024-09-04 18:34:45 +03:00
Avi Kivity	2b94ee981b	repair: row_level: row_level_repair::send_missing_rows_to_follower_nodes() The function itself is threaded, but the inner lambda is coroutinized.	2024-09-04 18:28:27 +03:00
Avi Kivity	c768448339	repair: row_level: row_level_repair::get_missing_rows_from_follower_nodes() The function itself is threaded, but the inner lambda is coroutinized.	2024-09-04 18:28:12 +03:00
Avi Kivity	d2f1b44487	repair: row_level: row_level_repair::negotiate_sync_boundary() The function itself is threaded, but the inner lambda is coroutinized.	2024-09-04 18:21:39 +03:00
Kefu Chai	0756520f82	sstable: coroutinize sstable::seal_sstable() for better readability. presumably, `sstable::seal_sstable()` is not on the critical path, and we don't need to worry about the overhead of using C++20 coroutine. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20410	2024-09-04 18:14:33 +03:00
Kefu Chai	88c5c3001a	compaction: refactor compaction_manager::can_proceed() instead of chaining the conditions with '&&', break them down. for two reasons: * for better readability: to group the conditions with the same purpose together * so we don't look up the table twice. it's an anti-pattern of using STL, and it could be confusing at first glance. this change is a cleanup, so it does not change the behavior. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20369	2024-09-04 18:12:29 +03:00
Avi Kivity	645e39e746	repair: row_level: coroutinize repair_put_row_diff_with_rpc_stream_process_op() Both the outer function and the inner lambda are coroutinized.	2024-09-04 18:10:43 +03:00
Avi Kivity	4c05d0b965	repair: row_level: coroutinize repair_meta::get_sync_boundary_handler()	2024-09-04 15:33:40 +03:00
Avi Kivity	eea011fad5	repair: row_level: coroutinize repair_meta::get_sync_boundary() Not really helping anything, but a coroutine is a safer platform for future changes in administrative APIs.	2024-09-04 15:31:57 +03:00
Avi Kivity	91b88df956	repair: row_level: coroutinize repair_meta::repair_set_estimated_partitions_handler()	2024-09-04 15:20:53 +03:00
Avi Kivity	b73194c9bf	repair: row_level: coroutinize repair_meta::repair_set_estimated_partitions() Not really helping anything, but a coroutine is a safer platform for future changes in administrative APIs.	2024-09-04 15:18:33 +03:00
Avi Kivity	a69fb626bd	repair: row_level: coroutinize repair_meta::repair_get_estimated_partitions_handler()	2024-09-04 15:17:42 +03:00
Avi Kivity	5cd8207ac7	repair: row_level: coroutinize repair_meta::repair_get_estimated_partitions() Not really helping anything, but a coroutine is a safer platform for future changes in administrative APIs.	2024-09-04 15:16:32 +03:00
Avi Kivity	e108f867a9	repair: row_level: coroutinize repair_meta::repair_row_level_stop_handler()	2024-09-04 15:15:42 +03:00
Avi Kivity	ffbb973063	repair: row_level: coroutinize repair_meta::repair_row_level_stop() Not really helping anything, but a coroutine is a safer platform for future changes in administrative APIs.	2024-09-04 15:14:08 +03:00
Avi Kivity	587b6fe400	repair: row_level: coroutinize repair_meta::repair_row_level_start_handler()	2024-09-04 15:12:49 +03:00
Avi Kivity	db7b1014ff	repair: row_level: coroutinize repair_meta::repair_row_level_start()	2024-09-04 15:10:45 +03:00
Avi Kivity	17b82265ae	repair: row_level: coroutinize repair_meta::get_combined_row_hash_handler()	2024-09-04 15:08:58 +03:00
Avi Kivity	bacbdde791	repair: row_level: coroutinize repair_meta::get_combined_row_hash()	2024-09-04 15:07:27 +03:00
Avi Kivity	8b8dc5092f	repair: row_level: coroutinize repair_meta::get_full_row_hashes_handler()	2024-09-04 15:05:28 +03:00
Avi Kivity	21e01990ff	repair: row_level: coroutinize repair_meta::get_full_row_hashes_with_rpc_stream() The when_all_succeed() call is changed to the safer coroutine::when_all(), which avoids the temporary futures.	2024-09-04 15:03:00 +03:00
Avi Kivity	572fbfde09	repair: row_level: coroutinize repair_meta::request_row_hashes()	2024-09-04 14:07:59 +03:00
Nadav Har'El	15f8046fcb	alternator ttl: fix use-after-free The Alternator TTL scanning code uses an object "scan_ranges_context" to hold the scanning context. One of the members of this object is a service::query_state, and that in turn holds a reference to a service::client_state. The existing constructor created a temporary client_state object and saved a reference to it - which can result in use after free as the temporary object is freed as soon as the constructor ends. The fix is to save a client_state in the scan_ranges_context object, instead of a temporary object. Fixes #19988 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20418	2024-09-03 22:15:18 +03:00
Pavel Emelyanov	c03b1e2827	test: Remove unused database argument from make_sstable_for_all_shards() helper Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20427	2024-09-03 21:36:28 +03:00
Calle Wilund	2695fefa81	commitlog/database: Make some commitlog options updatable + add feature listener Makes some commitlog options runtime updatable. Most important for this case, the usage of fragmented entries. Also adds a subscription in database on said feature, to possibly enable once cluster enables it.	2024-09-03 16:38:28 +00:00
Calle Wilund	238a0236e5	features/config: Add feature for fragmented commitlog entries Hides the functionality behind a cluster feature, i.e. postspones using it until an upgrade is complete etc. This to allow rolling back even with dirty nodes, at least until a cluster is commited. Feature can also be disabled by scylla option, just in case. This will lock it out of whole cluster, but this is probably good, because depending on off or on, certain schema/raft ops might fail or succeed (due to large mutations), and this should probably be equivalent across nodes.	2024-09-03 16:38:28 +00:00
Calle Wilund	9bf452c7a0	docs: Add entry on commitlog file format v4	2024-09-03 16:38:28 +00:00
Calle Wilund	ad595e4d6a	commitlog_test: Add more oversized cases Also adds some randomization to the tests.	2024-09-03 16:38:28 +00:00
Calle Wilund	1d5e509136	commitlog_replayer: Replay segments in order created Minimizes potential buffer usage for fragmented entries.	2024-09-03 16:38:28 +00:00
Calle Wilund	61ff9486fb	commitlog_replayer: Use replay state to support fragmented entries	2024-09-03 16:38:27 +00:00
Calle Wilund	7c16683184	commitlog_replayer: coroutinize partly	2024-09-03 16:38:27 +00:00
Calle Wilund	05bf2ae5d7	commitlog: Handle oversized entries Refs #18161 Yet another approach to dealing with large commitlog submissions. We handle oversize single mutation by adding yet another entry type: fragmented. In this case we only add a fragment (aha) of the data that needs storing into each entry, along with metadata to correlate and reconstruct the full entry on replay. Because these fragmented entries are spread over N segments, we also need to add references from the first segment in a chain to the subsequent ones. These are released once we clear the relevant cf_id count in the base. * This approach has the downside that due to how serialization etc works w.r.t. mutations, we need to create an intermediate buffer to hold the full serialized target entry. This is then incrementally written into entries of < max_mutation_size, successively requesting more segments. On replay, when encountering a fragment chain, the fragment is added to a "state", i.e. a mapping of currently processing frag chains. Once we've found all fragments and concatenated the buffers into a single fragmented one, we can issue a replay callback as usual. Note that a replay caller will need to create and provide such a state object. Old signature replay function remains for tests and such. This approach bumps the file format (docs to come). To ensure "atomicity" we both force syncronization, and should the whole op fail, we restore segment state (rewinding), thus discarding data all we wrote. v2: * Improve some bookeep, ensure we keep track of segments and flush properly, to get counter correct	2024-09-03 16:38:27 +00:00
Anna Stuchlik	35796306a7	doc: comment out redirections for pages under Features This commit temporarily disables redirections for all pages under Features that were moved with this PR: https://github.com/scylladb/scylladb/pull/20401 Redirections work for all versions. This means that pages in 6.1 are redirected to URLs that are not available yet (because 6.2 has not been released yet). The redirections are correct and should be enabled when 6.2 is released: I've created an issue to do it: https://github.com/scylladb/scylladb/issues/20428 Closes scylladb/scylladb#20429	2024-09-03 17:16:51 +02:00
Avi Kivity	6ddcf80d89	Merge 'Reuse sstable::test_env::reusable_sst() helper for pre-exsting sstables' from Pavel Emelyanov Tests that try to access sstables from test/resource/ typically sstable::load() it after object creation. There's reusable_sst() helper for that. This PR fixes one more caller that still goes longer route by doing sstable and loading it on its own. Closes scylladb/scylladb#20420 * github.com:scylladb/scylladb: test: Call reusable sst from ka_sst() helper test: Move sstable_open_config to reusable_sst()'s argument	2024-09-03 17:40:34 +03:00
Kamil Braun	504bf68ebb	raft_group0_client: remove `hold_read_apply_mutex` overload without `abort_source&` Ensure that every caller passes `abort_source&`.	2024-09-03 15:52:05 +02:00
Kamil Braun	79983723c8	storage_service: pass `_abort_source` to `hold_read_apply_mutex` There's no point waiting for this lock if `storage_service` is being aborted. In theory the lock, if held, should be eventually released by whatever is holding it during shutdown -- but if there is some cyclic reference between the services, and e.g. whatever holds the lock is stuck because of ongoing shutdown and would only be unstuck by `storage_service` getting stopped (which it can't because it's waiting on the lock), that would cause a shutdown deadlock. Better to be safe than sorry.	2024-09-03 15:52:05 +02:00
Kamil Braun	a7097fb985	group0_state_machine: pass `_abort_source` to `hold_read_apply_mutex` `transfer_snapshot` was already passing `_abort_source` when trying to take the lock but other member functions didn't.	2024-09-03 15:52:05 +02:00
Kamil Braun	a4d1065628	api: move `reload_raft_topology_state` implementation inside `storage_service` In later commit we'll want to access more `storage_service` internals in the API's implementation (namely, `_abort_source`) Also moving the implementation there allows making `service::topology_transition()` private again (it was made public in `992f1327d3` only for this API implementation)	2024-09-03 15:52:03 +02:00
Andrei Chekun	27e5fa149a	[test.py] Clean duplicated arg for test suite Arguments mode and run_id already set in the _prepare_pytest_params, so there is no need to set them one more time.	2024-09-03 14:41:57 +02:00
Andrei Chekun	8a9146ebda	[test.py] Enable allure for python test Enable allure adapter for all python tests. Add tag and parameters to the test to be able to distinguish them across modes and runs. Related: https://github.com/scylladb/qa-tasks/issues/1665	2024-09-03 14:41:57 +02:00
Łukasz Paszkowski	20a6296309	test: Add reversed query tests on simulated upgrade process Run the reversed queries on a 2-node cluster with CL=ALL with and without NATIVE_REVERSE_QUERIES feature flag. When the flag is enabled, the native reversed format is used, otherwise the legacy format. The NATIVE_REVERSE_QUERIES feature flag is suppressed with an error injection that simulates cluster upgrade process. Backport is not required. The patch adds additional upgrade tests for https://github.com/scylladb/scylladb/pull/18864 Closes scylladb/scylladb#20179	2024-09-03 14:45:08 +03:00
Pavel Emelyanov	0857b63259	Merge 'repair: row_level: coroutinize some slow-path functions' from Avi Kivity This series coroutinizes up some functions in repair/row_level.cc. This enhances readability and reduces bloat: ``` size build/release/repair/row_level.o.{before,after} text data bss dec hex filename 1650619 48 524 1651191 1931f7 build/release/repair/row_level.o.before 1604610 48 524 1605182 187e3e build/release/repair/row_level.o.after ``` 46kB of text were saved. Functions that only touch a single mutation fragment were not coroutinized to avoid adding a allocation in a fast path. In one case a function was split into a fast path and a slow path. Clean-up series, backport not needed. Closes scylladb/scylladb#20283 * github.com:scylladb/scylladb: repair: row_level: restore indentation repair: row_level: coroutinize repair_meta::get_full_row_hashes_sink_op() repair: row_level: coroutinize repair_meta::get_full_row_hashes_source_op() repair: row_level: coroutinize repair_get_full_row_hashes_with_rpc_stream_handler() repair: row_level: coroutinize repair_put_row_diff_with_rpc_stream_handler() repair: row_level: coroutinize repair_get_row_diff_with_rpc_stream_handler() repair: row_level: coroutinize repair_get_full_row_hashes_with_rpc_stream_process() repair: row_level: coroutinize repair_get_row_diff_with_rpc_stream_process_op_slow_path() repair: row_level: split repair_get_row_diff_with_rpc_stream_process_op() into fast and slow paths repair: row_level: coroutinize repair_meta::put_row_diff_handler() repair: row_level: coroutinize repair_meta::put_row_diff_sink_op() repair: row_level: coroutinize repair_meta::put_row_diff_source_op() repair: row_level: coroutinize repair_meta::put_row_diff() repair: row_level: coroutinize repair_meta::get_row_diff_handler() repair: row_level: coroutinize repair_meta::get_row_diff_sink_op() repair: row_level: coroutinize repair_meta::to_repair_rows_on_wire() repair: row_level: coroutinize repair_meta::do_apply_rows() repair: row_level: coroutinize repair_meta::copy_rows_from_working_row_buf_within_set_diff() repair: row_level: coroutinize repair_meta::copy_rows_from_working_row_buf() repair: row_level: coroutinize repair_meta::row_buf_csum() repair: row_level: coroutinize repair_meta::get_repairs_row_size() repair: row_level: coroutinize repair_meta::set_estimated_partitions() repair: row_level: coroutinize repair_meta::get_estimated_partitions() repair: row_level: coroutinize repair_meta::do_estimate_partitions_on_local_shard() repair: row_level: coroutinize repair_reader::close() repair: row_level: coroutinize repair_reader::end_of_stream() repair: row_level: coroutinize sink_source_for_repair::close() repair: row_level: coroutinize sink_source_for_repair::get_sink_source()	2024-09-03 14:41:22 +03:00
Nadav Har'El	dd030f8112	alternator: improve RBAC access denied error messages This patch address two requests made by reviewers of the original "Add CQL-based RBAC support to Alternator" series. Both requests were about the error messages produced when access is denied: 1. The error message is improved to use more proper English, and also to include the name of the role which was denied access. 2. The permission-check and error-message-formatting code is de-duplicated, using a common function verify_permission(). This de-duplication required moving the access-denied error path to throwing an exception instead of the previous exception-free implementation. However, it can be argued that this change is actually a good thing, because it makes the successful case, when access is allowed, faster. The de-duplicated code is shorter and simpler, and allowed changing the text of the error message in just one place. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20326	2024-09-03 14:39:30 +03:00
Kefu Chai	d26bb9ae30	sstables: correct the debugging message printed when removing temp dir in `372a4d1b79`, we introduced a change which was for debugging the logging message. but the logging message intended for printing the temp_dir not prints an `optional<int>`. this is both confusing, and more importantly, it hurts the debuggability. in this change, the related change is reverted. Fixes scylladb/scylladb#20408 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20409	2024-09-03 14:36:08 +03:00
Pavel Emelyanov	e4bc5470cf	test: Call reusable sst from ka_sst() helper The sstable_mutation_test wants to load pre-existing sstables from resouce/ subdir. For that there's reusable_sst() helper on env. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-03 14:01:28 +03:00
Pavel Emelyanov	e9980bd6dd	test: Move sstable_open_config to reusable_sst()'s argument So that callers are able to provide custom config in the future Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-03 14:00:59 +03:00
Laszlo Ersek	cd0819e3ed	docs/dev/docker-hub.md: refresh aio-max-nr calculation What we have today in "docs/dev/docker-hub.md" on "aio-max-nr" dates back to scylla commit `f4412029f4` ("docs/docker-hub.md: add quickstart section with --smp 1", 2020-09-22). Problems with the current language: - The "65K" claim as default value on non-production systems is wrong; "fs/aio.c" in Linux initializes "aio_max_nr" to 0x10000, which is 64K. - The section in question uses equal signs (=) incorrectly. The intent was probably to say "which means the same as", but that's not what equality means. - In the same section, the relational operator "<" is bogus. The available AIO count must be at least as high (>=) as the requested AIO count. - Clearer names should be used; adjust_max_networking_aio_io_control_blocks() in "src/core/reactor.cc" sets a great example: - "reactor::max_aio" should be called "storage_iocbs", - "detect_aio_poll" should be called "preempt_iocbs", - "reactor_backend_aio::max_polls" should be called "network_iocbs". - The specific value 10000 for the last one ("network_iocbs") is not correct in scylla's context. It is correct as the Seastar default, but scylla has used 50000 since commit `2cfc517874` ("main, test: adjust number of networking iocbs", 2021-07-18). Rewrite the section to address these problems. See also: - https://github.com/scylladb/scylladb/issues/5981 - https://github.com/scylladb/seastar/pull/2396 - https://github.com/scylladb/scylladb/pull/19921 Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-09-03 12:10:59 +02:00
Laszlo Ersek	15738d14ce	docs/dev/docker-hub.md: strip trailing whitespace Strip trailing whitespace. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-09-03 12:00:28 +02:00
Botond Dénes	2556e902b1	Update tools/jmx submodule * tools/jmx 89308b77...793452a9 (1): > dist: support building packages in Github Actions	2024-09-03 11:58:37 +03:00
Anna Stuchlik	5193d2d171	doc: remove the seeds-related questions from the FAQ This commit one of the series to remove the FAQ page by removing irrelevant/outdated entries or moving them to the forum. The question about seeds is irrelevant, not frequently asked, and covered in other sections of the docs. Also, it mentions versions that are no longer supported. Closes scylladb/scylladb#20403	2024-09-03 11:01:49 +03:00
Takuya ASADA	9d7fed40b5	install.sh: fix more incorrect permission on strict umask Even after `13caac7`, we still have more files incorrect permission, since we use "cp -r" and creating new file with redirect. To fix this, we need to replace "cp -r" with "cp -pr", and "chmod <perm>" on newly created files. Fixes #14383 Related #19775 Closes scylladb/scylladb#19786	2024-09-03 10:37:53 +03:00
Anna Stuchlik	360f7b3d33	doc: move Features to the top-level page This commit moves the Features page from the section for developers to the top level in the page tree. This involves: - Moving the source files to the features folder from the using-scylla folder. - Moving images into features/images folder. - Updating references to the moved resources. - Adding redirections to the moved pages. Closes scylladb/scylladb#20401	2024-09-03 07:24:33 +03:00
Kefu Chai	fb2ed20b42	.github: post a comment if "Fixes" policy is violated it's more visible than an "Error" in the action's detail message. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19271	2024-09-03 07:23:48 +03:00
Botond Dénes	8f31d3f1fc	Merge 'tools/nodetool: improve backup and restore commands' from Kefu Chai this change contains two improvements to "backup" and "restore" commands: - let them print task id - let them return 1 as the exist status code upon operation failure ---- these changes are improvements to the newly introduced commands, which are not in any LTS branches yet, so no need to backport. Closes scylladb/scylladb#20371 * github.com:scylladb/scylladb: tools/scylla-nodetool: return failure with exit code in backup/restore tools/scylla-nodetool: let backup/restore print task id	2024-09-02 16:40:55 +03:00
Takuya ASADA	59aedb38d0	locator: retry HTTP request to GCE/Azure metadata service Like we already do on EC2, implement retrying request to the metadata service on GCE and Azure. Closes #19817 Closes scylladb/scylladb#20189	2024-09-02 13:04:05 +03:00
Kefu Chai	e66e885e5b	tools/scylla-nodetool: return failure with exit code in backup/restore before this change, "backup" and "restore" commands always return 0 as their exist code no matter if the performed operation fails or not. inspired by the "task" commands of nodetool, let's return 1 with exit code if the operation fails. the tests are updated accordingly. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-02 15:12:26 +08:00
Kefu Chai	470c3e8535	tools/scylla-nodetool: let backup/restore print task id in `20fffcdc`, we added the "task wait" subcommand, so user is allowed to interact with a task with its task id. and in existing implementation of "backup" and "restore" command, if user does not pass `--nowait`, the command just exits without any output upon sending the request to scylladb. in this change, we print out the task_id if user does not pass `--nowait` command line option to "backup" or "restore" command. this allows user to follow up on the operation if necessary. the tests are updated accordingly. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-02 15:12:26 +08:00
Nadav Har'El	0b3890df46	test/cql-pytest: test RBAC auto-grant (and reproduce CDC bug) This patch adds functional testing for the role-based access control (RBAC) "auto-grant" feature, where if a user that is allowed to create a table, it also recieves full permissions over the table it just created. We also test permissions over new materialized views created by a user, and over CDC logs. The test for CDC logs reproduces an already suspected bug, #19798: A user may be allowed to create a table with CDC enabled, but then is not allowed to read the CDC log just created. The tests show that the other cases (base tables and views) do not have this bug, and the creating user does get appropriate permissions over the new table and views. In addition to testing auto-grant, the patch also includes tests for the opposite feature, "auto-revoke" - that permissions are removed when the table/view/cdc is deleted. If we forget to do that while implementing auto-grant, we risk that users may be able to use tables created by other users just because they used the same table _name_ earlier. It's important to have these auto-revoke tests together with the auto-grant tests that reproduce #19798 - so we don't forget this part when finally fixing #19798. Refs #19798. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#19845	2024-09-02 09:03:40 +03:00
Botond Dénes	52bed81a1e	Merge 'cql3: add option to not unify bind variables with the same name' from Avi Kivity Bind variables in CQL have two formats: positional (`?`) where a variable is referred to by its relative position in the statement, and named (`:var`), where the user is expected to supply a name->value mapping. In `19a6e69001` we identified the case where a named bind variable appears twice in a query, and collapsed it to a single entry in the statement metadata. Without this, a driver using the named variable syntax cannot disambiguate which variable is referred to. However, it turns out that users can use the positional call form even with the named variable syntax, by using the positional API of the driver. To support this use case, we add a configuration variable to disable the same-variable detection. Because the detection has to happen when the entire statement is visible, we have to supply the configuration to the parser. We call it the `dialect` and pass it from all callers. The alternative would be to add a pre-prepare call similar to fill_prepare_context that rewrites all expressions in a statement to deduplicate variables. A unit test is added. Fixes #15559 This may be useful to users transitioning from Cassandra, so merits a backport. Closes scylladb/scylladb#19493 * github.com:scylladb/scylladb: cql3: add option to not unify bind variables with the same name cql3: introduce dialect infrastructure cql3: prepared_statement_cache: drop cache key default constructor	2024-09-02 08:34:24 +03:00
Kefu Chai	28b5471c01	docs/dev/maintainer.md: fix formatting * in the "Backporting Seastar commits" section, there's a single quote instead of a backtick in this line, so fix it. * add backticks around `refresh-submodules.sh`, which is a filename. * correct the command line setting a git config option, because `git-config` does not support this command line syntax, ```console $ git config --global diff.conflictstyle = diff3 $ git config --global get diff.conflictstyle = $ git config --global diff.conflictstyle diff3 $ git config --global get diff.conflictstyle diff3 ``` quote from git-config(1) > ``` > git config set [<file-option>] [--type=<type>] [--all] [--value=<value>] [--fixed-value] <name> <value> > ``` * stop using the deprecated mode of the `git-config` command, and use subcommand instead. as git-config(1) puts: > git config <name> <value> [<value-pattern>] > Replaced by git config set [--value=<pattern>] <name> <value>. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20328	2024-09-01 22:24:01 +03:00
Yaniv Michael Kaul	2ebba9cd11	tools/toolchain/dbuild: prefer podman over docker Check if podman is available before docker. If it is, use it. Otherwise, check for docker. 1. Podman is better. It runs with fewer resources, and I've had display issues with Docker (output was not shown consistently) 2. 'which docker' works even when the docker service and socket are turned off. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Closes scylladb/scylladb#20342	2024-09-01 22:17:01 +03:00
David Garcia	c4da75e392	docs: run docs test on changing config params Triggers the "Build Docs" PR workflow whenever the `db/config.cc` or `db/config.h` files are edited. These files are used to produce documentation, and this change will help prevent the introduction of breaking changes to the documentation build when they are modified. Closes scylladb/scylladb#20347	2024-09-01 22:15:48 +03:00
Avi Kivity	0f4b05824e	Merge 'perf/perf_sstable: add {crawling,partitioned}_streaming modes' from Kefu Chai for testing the load performance of load_and_stream operation. Refs #19989 --- no need to backport. it adds two new tests to the existing `perf_sstable` tool for evaluating the load performance when performing the "load_and_streaming" operation. hence has no impact on the production. Closes scylladb/scylladb#20186 * github.com:scylladb/scylladb: perf/perf_sstable: add {crawling,partitioned}_streaming modes test/perf/perf_sstable: use switch-case when appropriate	2024-09-01 22:04:22 +03:00
Avi Kivity	7197d280b0	Merge 'scylla-gdb.py: lazy-evaluate the constants ' from Kefu Chai instead of evaluating the constants in-class, accessing them via a cached class property. it would be handy if we could source `scylla-gdb.py` in `.gdbinit`, but this script accesses some symbols which are not available without a file being debugged. what's why gdb fails to load the init script: ``` Traceback (most recent call last): File "/home/kefu/dev/scylladb/scylla-gdb.py", line 167, in <module> class intrusive_slist: File "/home/kefu/dev/scylladb/scylla-gdb.py", line 168, in intrusive_slist size_t = gdb.lookup_type('size_t') ^^^^^^^^^^^^^^^^^^^^^^^^^ gdb.error: No type named size_t. ``` so we have to `file path/to/scylla` and then `source scylla-gdb.py` every time when we debug scylla or a seastar application, instead of loading `scylla-gdb.py` in `.gdbinit`. the reason is that the script accesses the debug symbols like `gdb.lookup_type('size_t')` in-class. so when the python interpreter reads the script, it evaluates this statement, but at that moment, the debug symbols are not loaded, so `source scylla-gdb.py` fails in `.gdbinit`. in this change, we transform all these class variables to cached properties, so that they * are evaluated on-demand * are evaluated only once at most this addresses the pain at the expense of verbosity. --- this change intends to improve the developer's user experience, and has no impacts on product, so no need to backport. Closes scylladb/scylladb#20334 * github.com:scylladb/scylladb: test/scylla_gdb: test the .gdb init use case scylla-gdb.py: lazy-evaluate the constants	2024-09-01 20:00:53 +03:00
Pavel Emelyanov	7df43312ac	test: Remove sstable making helpers from table_for_tests All users of it have sstable_test_env at hand (in fact -- they call env method to get table_for_test). And since sstable_test_env already has a bunch of methods to create sstable, the table_for_test wrapper doesn't need to duplicate this code. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20360	2024-09-01 19:58:15 +03:00
Kefu Chai	bc2b7b47c8	build: cmake: add and use Scylla_CLANG_INLINE_THRESHOLD cmake parameter so that we can set this the parameter passed to `-inline-threshold` with `configure.py` when building with CMake. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20364	2024-09-01 19:56:02 +03:00
Kefu Chai	6970c502c9	dist: drop %pretrans section before this change, if user does not have `/bin/sh` around, when installing scylla packages, the script in `%pretrans" is executed, and fails due to missing `/bin/sh`. per https://docs.fedoraproject.org/en-US/packaging-guidelines/Scriptlets/#pretrans > Note that the %pretrans scriptlet will, in the particular case of > system installation, run before anything at all has been installed. > This implies that it cannot have any dependencies at all. For this > reason, %pretrans is best avoided, but if used it MUST (by necessity) > be written in Lua. See > https://rpm-software-management.github.io/rpm/manual/lua.html for more > information. but we were trying to warn users upgrading from scylla < 1.7.3, which was released 7 years ago at the time of writing. in this change, we drop the `%pretrans` section. hopefuly they will find their way out if they still exist. Fixes scylladb/scylladb#20321 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20365	2024-09-01 19:46:19 +03:00
Kefu Chai	a06e1c6545	scylla-housekeeping: use raw string to avoid using escape sequence before this change, when running `scylla-housekeeping`: ``` /opt/scylladb/scripts/libexec/scylla-housekeeping:122: SyntaxWarning: invalid escape sequence '\s' match = re.search(".http.?://repositories./scylladb/([^/\s]+)/./([^/\s]+)/scylladb-.", line) ``` we could have the warning above. because `\s` is not a valid escape sequence, but the Python interpreter accepts it as two separated characters of `\s` after complaining. but it's still annoying. so, let's use a raw string here. Refs scylladb/scylladb#20317 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20359	2024-09-01 18:59:23 +03:00
Kefu Chai	e431b90145	test/boost/view_build_test: include used header before this change, when building the test of `view_build_test` with clang-20, we can have following build failure: ``` FAILED: test/boost/CMakeFiles/view_build_test.dir/Debug/view_build_test.cc.o /home/kefu/.local/bin/clang++ -DBOOST_ALL_DYN_LINK -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TESTING_MAIN -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -isystem /home/kefu/dev/scylladb/abseil -isystem /home/kefu/dev/scylladb/build/rust -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb/build=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -fno-sanitize=vptr -MD -MT test/boost/CMakeFiles/view_build_test.dir/Debug/view_build_test.cc.o -MF test/boost/CMakeFiles/view_build_test.dir/Debug/view_build_test.cc.o.d -o test/boost/CMakeFiles/view_build_test.dir/Debug/view_build_test.cc.o -c /home/kefu/dev/scylladb/test/boost/view_build_test.cc /home/kefu/dev/scylladb/test/boost/view_build_test.cc:998:5: error: unknown type name 'simple_schema' 998 \| simple_schema ss; \| ^ ``` apparently, `simple_schema`'s declaration is not available in this translation unit. in this change * we include the header where `simple_schema` is defined, so that the build passes with clang-20. * also take this opportunity to reorder the header a little bit, so the testing headers are grouped together. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20367	2024-09-01 18:58:23 +03:00
Kefu Chai	753188c33d	test: include seastar/testing/random.hh when appropriate in a recent seastar change (644bb662), we do not include `seastar/testing/random.hh` in `seastar/testing/test_runner.hh` anymore, as the latter is not a facade of the former, and neither does it use the former. as a sequence, some tests which take the advantage of the included `seastar/testing/random.hh` do not build with the latest seastar: ``` FAILED: test/lib/CMakeFiles/test-lib.dir/key_utils.cc.o /usr/bin/clang++ -DBOOST_REGEX_DYN_LINK -DBOOST_REGEX_NO_LIB -DBOOST_UNIT_TEST_FRAMEWORK_DYN_LINK -DBOOST_UNIT_TEST_FRAMEWORK_NO_LIB -DDEVEL -DFMT_SHARED -DSCYLLA_BUILD_MODE=dev -DSCYLLA_ENABLE_ERROR_INJECTION -DSCYLLA_ENABLE_PREEMPTION_SOURCE -DSEASTAR_API_LEVEL=7 -DSEASTAR_ENABLE_ALLOC_FAILURE_INJECTION -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SSTRING -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -I/__w/scylladb/scylladb -I/__w/scylladb/scylladb/build/gen -I/__w/scylladb/scylladb/seastar/include -I/__w/scylladb/scylladb/build/seastar/gen/include -I/__w/scylladb/scylladb/build/seastar/gen/src -I/__w/scylladb/scylladb/build -isystem /__w/scylladb/scylladb/abseil -isystem /__w/scylladb/scylladb/build/rust -O2 -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-enum-constexpr-conversion -Wno-unused-parameter -ffile-prefix-map=/__w/scylladb/scylladb/build=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -Werror=unused-result -fstack-clash-protection -MD -MT test/lib/CMakeFiles/test-lib.dir/key_utils.cc.o -MF test/lib/CMakeFiles/test-lib.dir/key_utils.cc.o.d -o test/lib/CMakeFiles/test-lib.dir/key_utils.cc.o -c /__w/scylladb/scylladb/test/lib/key_utils.cc In file included from /__w/scylladb/scylladb/test/lib/key_utils.cc:11: /__w/scylladb/scylladb/test/lib/random_utils.hh:25:30: error: no member named 'local_random_engine' in namespace 'seastar::testing' 25 \| return seastar::testing::local_random_engine; \| ~~~~~~~~~~~~~~~~~~^ 1 error generated. ``` in this change, we include `seastar/testing/random.hh` when the random facility is used, so that they can be compiled with the latest seastar library. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20368	2024-09-01 18:57:07 +03:00
Kefu Chai	0104c7d371	tools/scylla-nodetool: s/vm.count()/vm.contains()/ under the hood, std::map::count() and std::map::contains() are nearly identical. both operations search for the given key witin the map. however, the former finds a equal range with the given key, and gets the distance between the disntance between the begin and the end of the range; while the later just searches with the given key. since scylla-nodetool is not a performance-critical application, the minor difference in efficiency between these two operations is unlikely to have a significant impact on its overall performance. while std::map::count() is generally suitable for our need, it might be beneficial to use a more appropriate API. in this change, we use std::map::contains() in the place of std::map::count() when checking for the existence of a paramter with given name. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20350	2024-09-01 18:39:00 +03:00
Avi Kivity	ddf344e4f1	Merge 'compaction: use structured binding and ranges library when appropriate' from Kefu Chai for better readability --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#20366 * github.com:scylladb/scylladb: compaction: use std::views::reverse when appropriate compaction: use structured binding when appropriate	2024-09-01 18:35:15 +03:00
Avi Kivity	ea8441dfa3	cql3: add option to not unify bind variables with the same name Bind variables in CQL have two formats: positional (`?`) where a variable is referred to by its relative position in the statement, and named (`:var`), where the user is expected to supply a name->value mapping. In `19a6e69001` we identified the case where a named bind variable appears twice in a query, and collapsed it to a single entry in the statement metadata. Without this, a driver using the named variable syntax cannot disambiguate which variable is referred to. However, it turns out that users can use the positional call form even with the named variable syntax, by using the positional API of the driver. To support this use case, we add a configuration variable to disable the same-variable detection. Because the detection has to happen when the entire statement is visible, we have to supply the configuration to the parser. We call it the `dialect` and pass it from all callers. The alternative would be to add a pre-prepare call similar to fill_prepare_context that rewrites all expressions in a statement to deduplicate variables. A unit test is added. Fixes #15559	2024-09-01 17:27:48 +03:00
Avi Kivity	60acfd8c08	docs: cql: document ZstdCompressor for CREATE TABLE Adjust the wording slightly to be less awkward. Closes scylladb/scylladb#20377	2024-09-01 14:28:09 +03:00
Kefu Chai	e53a9a99cd	compaction: use std::views::reverse when appropriate let's use the standard library when appropriate. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-01 08:44:01 +08:00
Kefu Chai	3801c079e2	compaction: use structured binding when appropriate for better readability Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-01 08:34:10 +08:00
Avi Kivity	61e6a77a99	repair: row_level: restore indentation	2024-08-30 23:00:59 +03:00
Avi Kivity	a35942e09a	repair: row_level: coroutinize repair_meta::get_full_row_hashes_sink_op() Extra care is needed for exception handling.	2024-08-30 22:55:16 +03:00
Avi Kivity	8e9ebd82fc	repair: row_level: coroutinize repair_meta::get_full_row_hashes_source_op()	2024-08-30 22:55:16 +03:00
Avi Kivity	f7d19e237d	repair: row_level: coroutinize repair_get_full_row_hashes_with_rpc_stream_handler() Both the handle_exception() and finally() blocks need some extra care.	2024-08-30 22:55:16 +03:00
Avi Kivity	bb8751f4b5	repair: row_level: coroutinize repair_put_row_diff_with_rpc_stream_handler() Both the handle_exception() and finally() blocks need some extra care.	2024-08-30 22:55:16 +03:00
Avi Kivity	7ba0642da2	repair: row_level: coroutinize repair_get_row_diff_with_rpc_stream_handler() Both the handle_exception() and finally() blocks need some extra care.	2024-08-30 22:55:16 +03:00
Avi Kivity	61bbf452c6	repair: row_level: coroutinize repair_get_full_row_hashes_with_rpc_stream_process()	2024-08-30 22:55:16 +03:00
Avi Kivity	01a578f608	repair: row_level: coroutinize repair_get_row_diff_with_rpc_stream_process_op_slow_path()	2024-08-30 22:55:16 +03:00
Avi Kivity	3733105f78	repair: row_level: split repair_get_row_diff_with_rpc_stream_process_op() into fast and slow paths This allows coroutinization of the slow path without affecting the fast path.	2024-08-30 22:55:16 +03:00
Avi Kivity	e17c3b71a8	repair: row_level: coroutinize repair_meta::put_row_diff_handler()	2024-08-30 22:55:16 +03:00
Avi Kivity	74ea2b9663	repair: row_level: coroutinize repair_meta::put_row_diff_sink_op() Exception handling is a bit awkward since can't co_await in a catch block.	2024-08-30 22:55:16 +03:00
Avi Kivity	e4362a5b7b	repair: row_level: coroutinize repair_meta::put_row_diff_source_op()	2024-08-30 22:55:16 +03:00
Avi Kivity	b998d69f09	repair: row_level: coroutinize repair_meta::put_row_diff()	2024-08-30 22:55:16 +03:00
Avi Kivity	3f2b5fe5dc	repair: row_level: coroutinize repair_meta::get_row_diff_handler()	2024-08-30 22:55:16 +03:00
Avi Kivity	cd63971501	repair: row_level: coroutinize repair_meta::get_row_diff_sink_op() Since sink.close() is called from an exception handler, some code movement is needed so it isn't co_awaited from a catch block.	2024-08-30 22:55:16 +03:00
Avi Kivity	3f28dec88c	repair: row_level: coroutinize repair_meta::to_repair_rows_on_wire() coroutine::maybe_yield() introduced to compensate for loss of stall-protected do_for_each()	2024-08-30 22:55:16 +03:00
Avi Kivity	1a84f1a73d	repair: row_level: coroutinize repair_meta::do_apply_rows() coroutine::maybe_yield() introduced to compensate for loss of stall-protected do_for_each()	2024-08-30 22:55:16 +03:00
Avi Kivity	7f15cc446f	repair: row_level: coroutinize repair_meta::copy_rows_from_working_row_buf_within_set_diff() coroutine::maybe_yield() introduced to compensate for loss of stall-protected do_for_each()	2024-08-30 22:55:16 +03:00
Avi Kivity	93ca202bd3	repair: row_level: coroutinize repair_meta::copy_rows_from_working_row_buf() coroutine::maybe_yield() introduced to compensate for loss of stall-protected do_for_each()	2024-08-30 22:55:15 +03:00
Avi Kivity	5f8895d908	repair: row_level: coroutinize repair_meta::row_buf_csum() coroutine::maybe_yield() introduced to compensate for loss of stall-protected do_for_each()	2024-08-30 22:55:15 +03:00
Avi Kivity	d1e45f2982	repair: row_level: coroutinize repair_meta::get_repairs_row_size() coroutine::maybe_yield() introduced to compensate for loss of stall-protected do_for_each()	2024-08-30 22:55:15 +03:00
Avi Kivity	0b1bf57d19	repair: row_level: coroutinize repair_meta::set_estimated_partitions()	2024-08-30 22:55:15 +03:00
Avi Kivity	aee078d8e5	repair: row_level: coroutinize repair_meta::get_estimated_partitions()	2024-08-30 22:55:15 +03:00
Avi Kivity	51534f60eb	repair: row_level: coroutinize repair_meta::do_estimate_partitions_on_local_shard()	2024-08-30 22:55:12 +03:00
Kamil Braun	e01cef01a6	Merge 'Ignore seed name resolution errors during the restart of a cluster member node.' from Sergey Zolotukhin All seeds hostname resolution errors will be ignored during a node restart in case the node had already joined a cluster. This will prevent restart errors if some seed names are not resolvable. Fixes scylladb/scylladb#14945 Closes scylladb/scylladb#20292 * github.com:scylladb/scylladb: Ignore seed name resolution errors on restart. Add a test for starting with a wrong seed.	2024-08-30 11:33:44 +02:00
Kamil Braun	292ef0d1f9	Merge 'Fix node replace with inter-dc encryption enabled.' from Gleb Natapov Currently if a coordinator and a node being replaced are in the same DC while inter-dc encryption is enabled (connections between nodes in the same DC should not be encrypted) the replace operation will fail. It fails because a coordinator uses non encrypted connection to push raft data to the new node, but the new node will not accept such connection until it knows which DC the coordinator belongs to and for that the raft data needs to be transferred. The series adds the test for this scenario and the fix for the chicken&egg problem above. The series (or at least the fix itself) needs to be backported because this is a serious regression. Fixes: scylladb/scylladb#19025 Closes scylladb/scylladb#20290 * github.com:scylladb/scylladb: topology coordinator: fix indentation after the last patch topology coordinator: do not add replacing node without a ring to topology test: add test for replace in clusters with encryption enabled test.py: add server encryption support to cluster manager .gitignore: fix pattern for resources to match only one specific directory	2024-08-30 11:29:05 +02:00
Kefu Chai	82fbe317ec	test/scylla_gdb: test the .gdb init use case before this change, we run all the tests in a single pytest session, with scylladb debug symbols loaded. but we want to test another use case, where the scylladb debug symbols are missing. in this change, * we do not check for the existence of debug symbols until necessary * add a mark named "without_scylla" * run the tests in two pytest sessions - one with "without_scylla" mark - one with "not without_scylla" mark * add a test which is marked with the "without_scylla" mark. the test verify that the scylla-gdb.py script can be loaded even without scylladb debug symbols. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-30 17:05:29 +08:00
Kefu Chai	7dd63c891f	scylla-gdb.py: lazy-evaluate the constants instead of evaluating the constants in-class, accessing them via a cached class property. it would be handy if we could source `scylla-gdb.py` in `.gdbinit`, but this script accesses some symbols which are not available with a file being debugged. so when gdb fails to load init script: ``` Traceback (most recent call last): File "/home/kefu/dev/scylladb/scylla-gdb.py", line 167, in <module> class intrusive_slist: File "/home/kefu/dev/scylladb/scylla-gdb.py", line 168, in intrusive_slist size_t = gdb.lookup_type('size_t') ^^^^^^^^^^^^^^^^^^^^^^^^^ gdb.error: No type named size_t. ``` so we have to `file path/to/scylla` and then `source scylla-gdb.py` every time when we debug scylla or a seastar application, instead of loading `scylla-gdb.py` in `.gdbinit`. the reason is that the script access the debug symbols like `gdb.lookup_type('size_t')` in-class. so when the python interpreter reads the script, it evaluates this statement, but at that moment, the debug symbols are not loaded, so `source scylla-gdb.py` fails in `.gdbinit`. in this change, we transform all these class variables to cached property, so that they * are evaluated on-demand * are evaluated only once at most this addresses the pain at the expense of verbosity. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-30 17:05:29 +08:00
Pavel Emelyanov	cec4d207f6	Merge 'repair: throw if batchlog manager isn't initialized' from Aleksandra Martyniuk repair_service::repair_flush_hints_batchlog_handler may access batchlog manager while it is uninitialized. Throw if batchlog manager isn't initialized. Fixes: #20236. Needs backport to 6.0 and 6.1 as they suffer from the uninitialized bm access. Closes scylladb/scylladb#20251 * github.com:scylladb/scylladb: test: add test to ensure repair won't fail with uninitialized bm repair: throw if batchlog manager isn't initialized	2024-08-30 11:37:24 +03:00
Anna Stuchlik	4471c80bdc	doc: add the 6.1-to-6.2 upgrade guide This commit replaces the 6.0-to-6.1 upgrade guide with the 6.1-to-6.2 upgrade guide. The new guide is a template that covers the basic procedure. If any 6.2-specific updates are required, they will have to be added along with development. Closes scylladb/scylladb#20178	2024-08-30 10:10:45 +03:00
Piotr Dulikowski	c05be27e4a	Merge 'db/hints: Move the code for writing hints to a separate function' from Dawid Mędrek In scylladb/scylladb@7301a96, in the function `hint_endpoint_manager::store_hint()`, we transformed the lambda passed to `seastar::with_gate()` to a coroutine lambda to improve the readability. However, there was a subtle problem related to lifetimes of the captures that needed to be addressed: * Since we started `co_await`ing in the lambda, the captures were at risk of being destructed too soon. The usual solution is to wrap a coroutine lambda within a `seastar::coroutine::lambda` object and rely on the extended lifetime enforced by the semantics of the language. See `docs/dev/lambda-coroutine-fiasco.md` for more context. * However, since we don't immediately `co_await` the future returned by `with_gate()`, we cannot rely on the extended lifetime provided by the wrapper. The document linked in the previous bullet point suggests keeping the passed coroutine lambda as a variable and pass it as a reference to `with_gate()`. However, that's not feasible either because we discard the returned future and the function returns almost instantly -- destructing every local object, which would encompass the lambda too. The solution used in the commit was to move captures of the lambda into the lambda's body. That helped because Seastar's backend is responsible for keeping all of the local variables alive until the lambda finishes its execution. However, we didn't move all of the captures into the lambda -- the missing one was the `this` pointer that was implicitly used in the lambda. Address sanitiser hasn't reported any bugs related to the pointer yet, but the bug is most likely there. In this commit, we transform the lambda's body into a new member function and only call it from the lambda. This way, we don't need to care about the lifetimes of the captures because Seastar ensures that the function's arguments stay alive until the coroutine finishes. Choosing this solution instead of assigning `this` to a pointer variable inside the lambda's body and using it to refer to the object's members has actual benefit: it's not possible to accidentally forget to refer to a member of the object via the pointer; it also makes the code less awkward. Fixes scylladb/scylladb#20306 Closes scylladb/scylladb#20258 * github.com:scylladb/scylladb: db/hints: Fix indentation in `do_store_hint()` db/hints: Move code for writing hints to separate function	2024-08-30 09:09:02 +02:00
Avi Kivity	bbcfd47bf5	doc: nodetool: toppartitions: document --samplers and --capacity In particular --capacity is critical for obtaining accurate measurements. Closes scylladb/scylladb#20192	2024-08-30 10:07:54 +03:00
Botond Dénes	9f9346fc59	Merge 'nodetool: tasks: add nodetool commands to track task manager tasks' from Aleksandra Martyniuk Add nodetool commands to manage task manager tasks: - tasks abort - aborts the task - tasks list - lists all tasks in the module - tasks modules - lists all modules - tasks set-ttl - sets task ttl - tasks status - gets status of the task - tasks tree - gets statuses of the task and all its desendent's - tasks ttl - gets task ttl - tasks wait - waits for the task and gets its status Fixes: https://github.com/scylladb/scylladb/issues/19201. Closes scylladb/scylladb#19614 * github.com:scylladb/scylladb: test: nodetool: add tests for tasks commands nodetool: tasks: add nodetool commands to track task manager tasks api: task_manager: return status 403 if a task is not abortable api: task_manager: return none instead of empty task id api: task_manager: add timeout to wait_task api: task_manager: add operation to get ttl nodetool: add suboperations support nodetool: change operations_with_func type nodetool: prepare operation related classes for suboperations	2024-08-30 07:37:37 +03:00
Avi Kivity	d69bf4f010	cql3: introduce dialect infrastructure A dialect is a different way to interpret the same CQL statement. Examples: - how duplicate bind variable names are handled (later in this series) - whether `column = NULL` in LWT can return true (as is now) or whether it always returns NULL (as in SQL) Currently, dialect is an empty structure and will be filled in later. It is passed to query_processor methods that also accept a CQL string, and from there to the parser. It is part of the prepared statement cache key, so that if the dialect is changed online, previous parses of the statement are ignored and the statement is prepared again. The patch is careful to pick up the dialect at the entry point (e.g. CQL protocol server) so that the dialect doesn't change while a statement is parsed, prepared, and cached.	2024-08-29 21:19:23 +03:00
Avi Kivity	f9322799af	cql3: prepared_statement_cache: drop cache key default constructor It's unnecessary, and interferes with the following patch where we change the cache key type.	2024-08-29 21:07:00 +03:00
Avi Kivity	67b24859bc	Merge 'generic_server: convert connection tracking to seastar::gate' from Laszlo Ersek ~~~ generic_server: convert connection tracking to seastar::gate If we call server::stop() right after "server" construction, it hangs: With the server never listening (never accepting connections and never serving connections), nothing ever calls server::maybe_stop(). Consequently, co_await _all_connections_stopped.get_future(); at the end of server::stop() deadlocks. Such a server::stop() call does occur in controller::do_start_server() [transport/controller.cc], when - cserver->start() (sharded<cql_server>::start()) constructs a "server"-derived object, - start_listening_on_tcp_sockets() throws an exception before reaching listen_on_all_shards() (for example because it fails to set up client encryption -- certificate file is inaccessible etc.), - the "deferred_action" cserver->stop().get(); is invoked during cleanup. (The cserver->stop() call exposing the connection tracking problem dates back to commit `ae4d5a60ca` ("transport::controller: Shut down distributed object on startup exception", 2020-11-25), and it's been triggerable through the above code path since commit `6b178f9a4a` ("transport/controller: split configuring sockets into separate functions", 2024-02-05).) Tracking live connections and connection acceptances seems like a good fit for "seastar::gate", so rewrite the tracking with that. "seastar::gate" can be closed (and the returned future can be waited for) without anyone ever having entered the gate. NOTE: this change makes it quite clear that neither server::stop() nor server::shutdown() must be called multiple times. The permitted sequences are: - server::shutdown() + server::stop() - or just server::stop(). Fixes #10305 Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> ~~~ Fixes #10305. I think we might want to backport this -- it fixes a hang-on-misconfiguration which affects `scylla-6.1.0-0.20240804.abbf0b24a60c.x86_64` minimally. Basically every release that contains commit `ae4d5a60ca` has a theoretical chance for the hang, and every release that contains commit `6b178f9a4a` has a practical chance for the hang. Focusing on the more practical symptom (i.e., releases containing commit `6b178f9a4a`), `git tag --contains 6b178f9a4a90` gives us (ignoring candidates and release candidates): - scylla-6.0.0 - scylla-6.0.1 - scylla-6.0.2 - scylla-6.1.0 Closes scylladb/scylladb#20212 * github.com:scylladb/scylladb: generic_server: make server::stop() idempotent generic_server: coroutinize server::shutdown() generic_server: make server::shutdown() idempotent test/generic_server: add test case configure, cmake: sort the lists of boost unit tests generic_server: convert connection tracking to seastar::gate	2024-08-29 19:45:48 +03:00
Laszlo Ersek	db44000f8d	Update seastar submodule * seastar 83e6cdfd...ec5da7a6 (1): > reactor, linux-aio: advise users in more detail on setting aio-max-nr Fixes #5981 Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#20307	2024-08-29 19:42:02 +03:00
Raphael S. Carvalho	26facd807e	storage_service: avoid processing same table unnecessarily in split monitor If there's a token metadata for a given table, and it is in split mode, it will be registered such that split monitor can look at it, for example, to start split work, or do nothing if table completed it. during topology change, e.g. drain, split is stalled since it cannot take over the state machine. It was noticed that the log is being spammed with a message saying the table completed split work, since every tablet metadata update, means waking up the monitor on behalf of a table. So it makes sense to demote the logging level to debug. That persists until drain completes and split can finally complete. Another thing that was noticed is that during drain, a table can be submitted for processing faster than the monitor can handle, so the candidate queue may end up with multiple duplicated entries for same table, which means unnecessary work. That is fixed by using a sequenced set, which keeps the current FIFO behavior. Fixes #20339. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Closes scylladb/scylladb#20029	2024-08-29 19:38:43 +03:00
Aleksandra Martyniuk	1f46cad5de	test: nodetool: add tests for tasks commands	2024-08-29 17:37:13 +02:00
Aleksandra Martyniuk	20fffcdcf5	nodetool: tasks: add nodetool commands to track task manager tasks	2024-08-29 17:37:12 +02:00
Avi Kivity	7da3314deb	Merge 'Integrated restore' from Ernest Zaslavsky Handed over from https://github.com/scylladb/scylladb/pull/20149 This adds minimal implementation of the start-restore API call. The method starts a task that runs load-and-stream functionality against sstables from S3 bucket. Arguments are: ``` endpoint -- the ID in object_store.yaml config file bucket -- the target bucket to get objects from keyspace -- the keyspace to work on table -- the table to work on snapshot -- the name of the snapshot from which the backup was taken ``` The task runs in the background, its task_id is returned from the method once it's spawned and it should be used via /task_manager API to track the task execution and completion. Remote sstables components are scanned as if they were placed in local upload/ directory. Then colelcted sstables are fed into load-and-stream. This branch has https://github.com/scylladb/scylladb/pull/19890 (Integrated backup), https://github.com/scylladb/scylladb/pull/20120 (S3 lister) and few more minor PRs merged in. The restore branch itself starts with [utils: Introduce abstract (directory) lister](`29c867b54d`) commit. refs: https://github.com/scylladb/scylladb/issues/18392 Closes scylladb/scylladb#20305 * github.com:scylladb/scylladb: tools/scylla-nodetool: add restore integration test/object_store: Add simple restore test test/object_store: Generalize prepare_snapshot_for_backup() code: Introduce restore API method sstable_loader: Add sstables::storage_manager dependency sstable_loader: Maintain task manager module sstable_loader: Out-line constructor distributed_loader: Split get_sstables_from_upload_dir() sstables/storage: Compose uploaded sstable path simpler sstable_directory: Prepare FS lister to scan files on S3 sstable_directory: Parse sstable component without full path s3-client: Add support for lister::filter utils: Introduce abstract (directory) lister	2024-08-29 18:25:30 +03:00
Kamil Braun	9574c399ce	Merge 'add support for zero-token nodes' from Patryk Jędrzejczak We revive the `join_ring` option. We support it only in the Raft-based topology, as we plan to remove the gossip-based topology when we fix the last blocker - the implementation of the manual recovery tool. In the Raft-based topology, a node can be assigned tokens only once when it joins the cluster. Hence, we disallow joining the ring later, which is possible in Cassandra. The main idea behind the solution is simple. We make the unsupported special case of zero tokens a supported normal case. Nodes with zero tokens assigned are called "zero-token nodes" from now on. From the topology point of view, zero-token nodes are the same as token-owning nodes. They can be in the same states, etc. From the data point of view, they are different. They are not members of the token ring, so they are not present in `token_metadata::_normal_token_owners`. Hence, they are ignored in all non-local replication strategies. The tablet load balancer also ignores them. Zero-token nodes can be used as coordinator-only nodes, just like in Cassandra. They can handle requests just like token-owning nodes. The main motivation behind zero-token nodes is that they can prevent the Raft majority loss efficiently. Zero-token nodes are group 0 voters, but they can run on much weaker and cheaper machines because they do not replicate data and handle client requests by default (drivers ignore them). For example, if there are two DCs, one with 4 nodes and one with 5 nodes, if we add a DC with 2 zero-token nodes, every DC will contain less than half of the nodes, so we won't lose the majority when any DC dies. Another way of preventing the Raft majority loss is changing the voter set, which is tracked by scylladb/scylladb#18793. That approach can be used together with zero-token nodes. In the example above, if we choose equal numbers of voters in both DCs, then a DC with one zero-token node will be sufficient. However, in the typical setup of 2 DCs with the same number of nodes it is enough to add a DC with only one zero-token node without changing the voter set. Zero-token nodes could also be used as load balancers in the Alternator. Additionally, this PR fixes scylladb/scylladb#11087, which turned out to be a blocker. This PR introduced a new feature. There is no need to backport it. Fixes scylladb/scylladb#6527 Fixes scylladb/scylladb#11087 Fixes scylladb/scylladb#15360 Closes scylladb/scylladb#19684 * github.com:scylladb/scylladb: docs: raft: document using zero-token nodes to prevent majority loss test: test recovery mode in the presence of zero-token nodes test: topology: util.py: add cqls parameter to check_system_topology_and_cdc_generations_v3_consistency test: topology: util.py: accept zero tokens in check_system_topology_and_cdc_generations_v3_consistency treewide: support zero-token nodes in the recovery mode storage_proxy: make TRUNCATE work locally for local tables test: topology: util.py: document that check_token_ring_and_group0_consistency fails with zero-token nodes test: test zero-token nodes test: test_topology_ops: move helpers to topology/util.py feature_service: introduce the ZERO_TOKEN_NODES feature storage_service: rename join_token_ring to join_topology storage_service: raft_topology_cmd_handler: improve warnings topology_coordinator: fix indentation after the previous patch treewide: introduce support for zero-token nodes in Raft topology system_keyspace: load_topology_state: remove assertion impossible to hit treewide: distinguish all nodes from all token owners gossip topology: make a replacing node remove the replaced node from topology locator: topology: add_or_update_endpoint: use none as the default node state test: boost: tablets tests: ensure all nodes are normal token owners token_metadata: rename get_all_endpoints and get_all_ips network_topology_strategy: reallocate_tablets: remove unused dc_rack_nodes virtual_tables: cluster_status_table: execute: set dc regardless of the token ownership	2024-08-29 16:26:21 +02:00
Gleb Natapov	32a59ba98f	topology coordinator: fix indentation after the last patch	2024-08-29 17:14:09 +03:00
Gleb Natapov	17f4a151ce	topology coordinator: do not add replacing node without a ring to topology When only inter dc encryption is enabled a non encrypted connection between two nodes is allowed only if both nodes are in the same dc. If a nodes that initiates the connection knows that dst is in the same dc and hence use non encrypted connection, but the dst not yet knows the topology of the src such connection will not be allowed since dst cannot guaranty that dst is in the same dc. Currently, when topology coordinator is used, a replacing node will appear in the coordinator's topology immediately after it is added to the group0. The coordinator will try to send raft message to the new node and (assuming only inter dc encryption is enabled and replacing node and the coordinator are in the same dc) it will try to open regular, non encrypted, connection to it. But the replacing node will not have the coordinator in it's topology yet (it needs to sync the raft state for that). so it will reject such connection. To solve the problem the patch does not add a replacing node that was just added to group0 to the topology. It will be added later, when tokens will be assigned to it. At this point a replacing node will already make sure that its topology state is up-to-date (since it will execute a raft barrier in join_node_response_params handler) and it knows coordinator's topology. This aligns replace behaviour with bootstrap since bootstrap also does not add a node without a ring to the topology. The patch effectively reverts `b8ee8911ca` Fixes: scylladb/scylladb#19025	2024-08-29 17:14:09 +03:00
Gleb Natapov	2f1b1fd45e	test: add test for replace in clusters with encryption enabled	2024-08-29 17:14:09 +03:00
Gleb Natapov	b98282a976	test.py: add server encryption support to cluster manager	2024-08-29 17:14:09 +03:00
Gleb Natapov	84757a4ed3	.gitignore: fix pattern for resources to match only one specific directory	2024-08-29 17:13:58 +03:00
Dawid Medrek	d459cf91eb	db/hints: Fix indentation in `do_store_hint()`	2024-08-29 14:47:08 +02:00
Dawid Medrek	75ce6943d0	db/hints: Move code for writing hints to separate function In scylladb/scylladb@7301a96, in the function `hint_endpoint_manager::store_hint()`, we transformed the lambda passed to `seastar::with_gate()` to a coroutine lambda to improve the readability. However, there was a subtle problem related to lifetimes of the captures that needed to be addressed: * Since we started `co_await`ing in the lambda, the captures were at risk of being destructed too soon. The usual solution is to wrap a coroutine lambda within a `seastar::coroutine::lambda` object and rely on the extended lifetime enforced by the semantics of the language. See `docs/dev/lambda-coroutine-fiasco.md` for more context. * However, since we don't immediately `co_await` the future returned by `with_gate()`, we cannot rely on the extended lifetime provided by the wrapper. The document linked in the previous bullet point suggests keeping the passed coroutine lambda as a variable and pass it as a reference to `with_gate()`. However, that's not feasible either because we discard the returned future and the function returns almost instantly -- destructing every local object, which would encompass the lambda too. The solution used in the commit was to move captures of the lambda into the lambda's body. That helped because Seastar's backend is responsible for keeping all of the local variables alive until the lambda finishes its execution. However, we didn't move all of the captures into the lambda -- the missing one was the `this` pointer that was implicitly used in the lambda. Address sanitiser hasn't reported any bugs related to the pointer yet, but the bug is most likely there. In this commit, we transform the lambda's body into a new member function and only call it from the lambda. This way, we don't need to care about the lifetimes of the captures because Seastar ensures that the function's arguments stay alive until the coroutine finishes. Choosing this solution instead of assigning `this` to a pointer variable inside the lambda's body and using it to refer to the object's members has actual benefit: it's not possible to accidentally forget to refer to a member of the object via the pointer; it also makes the code less awkward.	2024-08-29 14:47:02 +02:00
Aleksandra Martyniuk	627fc46ca7	api: task_manager: return status 403 if a task is not abortable	2024-08-29 13:53:40 +02:00
Aleksandra Martyniuk	10ab60f32b	api: task_manager: return none instead of empty task id If a user requests a status of a task that does not have a parent, show "none" instead of an empty parent_id.	2024-08-29 13:53:40 +02:00
Aleksandra Martyniuk	5bcff4d544	api: task_manager: add timeout to wait_task	2024-08-29 13:53:40 +02:00
Aleksandra Martyniuk	3d78172328	api: task_manager: add operation to get ttl	2024-08-29 13:53:39 +02:00
Aleksandra Martyniuk	fb160afaf6	nodetool: add suboperations support Modify nodetool methods so that it support suboperations.	2024-08-29 13:53:39 +02:00
Aleksandra Martyniuk	4b96f9abb9	nodetool: change operations_with_func type Change the type of operations_with_func so that they can contain suboperations.	2024-08-29 13:53:39 +02:00
Aleksandra Martyniuk	c6f8a0116a	nodetool: prepare operation related classes for suboperations Modify operation and add operation_action class so that information about suboperations is stored. It's a preparation for adding suboperations support to nodetool.	2024-08-29 13:53:39 +02:00
Kefu Chai	dbb056f4f7	build: cmake: point -ffile-prefix-map to build directory before this change, we included `-ffile-prefix-map=${CMAKE_SOURCE_DIR}=.` in cflags when building the tree with CMake, but this was wrong. as the "." directory is the build directory used by CMake. and this directory is specified by the `-B` option when generating the building system. if `configure.py --use-cmake` is used to build the tree, the build directory would be "build". so this option instructs the compiler to replace the directory of source file in the debug symbols and in `__FILE__` at compile time. but, in a typical workspace, for instance, `build/main.cc` does not exist. the reason why this does not apply to CMake but applies to the rules generated by `configure.py` is that, `configure.py` puts the generated `build.ninja` right under the top source directory, so `.` is correct and it helps to create reproducible builds. because this practically erases the path prefixes in the build output. while CMake puts it under the specified build directory, replacing the source directory with the build directory with the file prefix map is just wrong. there are two options to address this problem: * stop passing this option. but this would lead to non-reproducible builds. as we would encode the build directory in the "scylla" executable. if a developer needs to rebuild an executable for debugging a coredump generated in production, he/she would have to either build the tree in the same directory as our CI does. or, he/she has to pass `-ffile-prefix-map=...` to map the local build directory to the one used by CI. this is not convenient. * instead of using `${CMAKE_SOURCE_DIR}=.`, add `${CMAKE_BINARY_DIR}=.`. this erases the build directory in the outputs, but preserves the debuggability. so we pick the second solution. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20329	2024-08-29 12:28:11 +03:00
Patryk Jędrzejczak	c192a9ee3b	docs: raft: document using zero-token nodes to prevent majority loss	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	e027ffdffc	test: test recovery mode in the presence of zero-token nodes We modify existing tests to verify that the recovery mode works correctly in the presence of zero-token nodes. In `test_topology_recovery_basic`, we test the case when a zero-token node is live. In particular, we test that the gossip-based restart of such a node works. In `test_topology_recovery_after_majority_loss`, we test the case when zero-token nodes are unrecoverable. In particular, we test that the gossip-based removenode of such nodes works. Since zero-token nodes are ignored by the Python driver if it also connects to other nodes, we use different CQL sessions for a zero-token node in `test_topology_recovery_basic`.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	fb1e060c4c	test: topology: util.py: add cqls parameter to check_system_topology_and_cdc_generations_v3_consistency In the following commit, we modify `test_topology_recovery_basic` to test the recovery mode in the presence of live zero-token nodes. Unfortunately, it requires a bit ugly workaround. Zero-token nodes are ignored by the Python driver if it also connects to other nodes because of empty tokens in the `system.peers` table. In that test, we must connect to a zero-token node to enter the recovery mode and purge the Raft data. Hence, we use different CQL sessions for different nodes. In the future, we may change the Python driver behavior and revert this workaround. Moreover, the recovery tests will be removed or significantly changed when we implement the manual recovery tool. Therefore, we shouldn't worry about this workaround too much.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	54905fc179	test: topology: util.py: accept zero tokens in check_system_topology_and_cdc_generations_v3_consistency Before we use `check_system_topology_and_cdc_generations_v3_consistency` in a test with a zero-token node, we must ensure it doesn't fail because of zero tokens in a row of the `system.topology` table.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	02bb70da19	treewide: support zero-token nodes in the recovery mode Before we implement the manual recovery tool, we must support zero-token nodes in the recovery mode. This means that two topology operations involving zero-token nodes must work in the gossip-based topology: - removing a dead zero-token node, - restarting a live zero-token node. We make changes necessary to make them work in this patch.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	87b415efdc	storage_proxy: make TRUNCATE work locally for local tables In on of the following patches, we implement support for zero-token nodes in the recovery mode. To achieve this, we need to be able to purge all Raft data on live zero-token nodes by using TRUNCATE. Currently, TRUNCATE works the same for all replication strategies - it is performed on all token owners. However, zero-token nodes are not token owners, so TRUNCATE would ignore them. Since zero-token nodes store only local tables, fixing scylladb/scylladb#11087 is the perfect solution for the issue with zero-token nodes. We do it in this patch. Fixes scylladb/scylladb#11087	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	21c8409fa4	test: topology: util.py: document that check_token_ring_and_group0_consistency fails with zero-token nodes	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	95e14ae44b	test: test zero-token nodes We add tests to verify the basic properties of zero-token nodes. `test_zero_token_nodes_no_replication` and `test_not_enough_token_owners` are more or less deterministic tests. Running them only in the dev mode is sufficient. `test_zero_token_nodes_topology_ops` is quite slow, as expected, considering parameterization and the number of topology operations. In the future we can think of making it faster or skipping in the debug mode. For now, our priority is to test zero-token nodes thoroughly.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	d43d67c525	test: test_topology_ops: move helpers to topology/util.py In one of the following patches, we reuse the helper functions from `test_topology_ops` in a new test, so we move them to `util.py`. Also, we add the `cl` parameter to `start_writes`, as the new test will use `cl=2`.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	574c252391	feature_service: introduce the ZERO_TOKEN_NODES feature Zero-token nodes must be supported by all nodes in the cluster. Otherwise, the non-supporting nodes would crash on some assertion that assumes only token-owing normal nodes make sense. Hence, we introduce the ZERO_TOKEN_NODES cluster feature. Zero-token nodes refuse to boot if it is not supported. I tested this patch manually. First, I booted a node built in the previous patch. Then, I tried to add a zero-token node built in this patch. It refused to boot as expected.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	c25eefe217	storage_service: rename join_token_ring to join_topology After introducing zero-token nodes that call join_token_ring but do not join the ring, the join_token_ring name does not make much sense.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	9937cf3a24	storage_service: raft_topology_cmd_handler: improve warnings	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	3ce936da7b	topology_coordinator: fix indentation after the previous patch	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	22d907e721	treewide: introduce support for zero-token nodes in Raft topology We revive the `join_ring` option. We support it only in the Raft-based topology, as we plan to remove the gossip-based topology when we fix the last blocker - the implementation of the manual recovery tool. In the Raft-based topology, a node can be assigned tokens only once when it joins the cluster. Hence, we disallow joining the ring later, which is possible in Cassandra. The main idea behind the solution is simple. We make the unsupported special case of zero tokens a supported normal case. Nodes with zero tokens assigned are called "zero-token nodes" from now on. From the topology point of view, zero-token nodes are the same as token-owning nodes. They can be in the same states, etc. From the data point of view, they are different. They are not members of the token ring, so they are not present in `token_metadata::_normal_token_owners`. Hence, they are ignored in all non-local replication strategies. The tablet load balancer also ignores them. Topology operations involving zero-token nodes are simplified: - `add` and `replace` finish in the `join_group0` state, so creating a new CDC generation and streaming are skipped, - `removenode` and `decommission` skip streaming, - `rebuild` does not even contact the topology coordinator as there is nothing to rebuild, Also, if the topology operation involves a token-owning node, zero-token nodes are ignored in streaming. Zero-token nodes can be used as coordinator-only nodes, just like in Cassandra. They can handle requests just like token-owning nodes. The main motivation behind zero-token nodes is that they can prevent the Raft majority loss efficiently. Zero-token nodes are group 0 voters, but they can run on much weaker and cheaper machines because they do not replicate data and handle client requests by default (drivers ignore them). For example, if there are two DCs, one with 4 nodes and one with 5 nodes, if we add a DC with 2 zero-token nodes, every DC will contain less than half of the nodes, so we won't lose the majority when any DC dies. Another way of preventing the Raft majority loss is changing the voter set, which is tracked by scylladb/scylladb#18793. That approach can be used together with zero-token nodes. In the example above, if we choose equal numbers of voters in both DCs, then a DC with one zero-token node will be sufficient. However, in the typical setup of 2 DCs with the same number of nodes it is enough to add a DC with only one zero-token node without changing the voter set. Zero-token nodes could also be used as load balancers in the Alternator.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	ba016c9af7	system_keyspace: load_topology_state: remove assertion impossible to hit We store tokens in a non-frozen set, which doesn't distinguish an empty set from no value. Hence, hitting this assertion is impossible.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	ed55261650	treewide: distinguish all nodes from all token owners In one of the following patches, we introduce support for zero-token nodes. From that point, getting all nodes and getting all token owners isn't equivalent. In this patch, we ensure that we consider only token owners when we want to consider only token owners (for example, in the replication logic), and we consider all nodes when we want to consider all nodes (for example, in the topology logic). The main purpose of this patch is to make the PR introducing zero-token nodes easier to review. The patch that introduces zero-token nodes is already complicated. We don't want trivial changes from this patch to make noise there. This patch introduces changes needed for zero-token nodes only in the Raft-based topology and in the recovery mode. Zero-token nodes are unsupported in the gossip-based topology outside recovery. Some functions added to `token_metadata` and `topology` are inefficient because they compute a new data structure in every call. They are never called in the hot path, so it's not a serious problem. Nevertheless, we should improve it somehow. Note that it's not obvious how to do it because we don't want to make `token_metadata` store topology-related data. Similarly, we don't want to make `topology` store token-related data. We can think of an improvement in a follow-up. We don't remove unused `topology::get_datacenter_rack_nodes` and `topology::get_datacenter_nodes`. These function can be useful in the future. Also, `topology::_dc_nodes` is used internally in `topology`.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	2d9575d6a9	gossip topology: make a replacing node remove the replaced node from topology In the following patch, we change the gossiper to work the same for zero-token nodes and token-owning nodes. We replace occurrences of `is_normal_token_owner` with topology-based conditions. We want to rely on the invariant that token-owning nodes own tokens if and only if they are in the normal or leaving state. However, this invariant is broken by a replacing node because it does not remove the replaced node from topology. Hence, after joining, the replacing node has topology with a node that is not a token owner anymore but is in a leaving state (`being_replaced`). We fix it to prevent the following patch from introducing a regression.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	c7016dedb3	locator: topology: add_or_update_endpoint: use none as the default node state In one of the following patches, we change the gossiper to work the same for zero-token nodes and token-owning nodes. We replace occurrences of `is_normal_token_owner` with topology-based conditions. We want to rely on the invariant that token-owning nodes own tokens if and only if they are in the normal or leaving state. However, this invariant can be broken in the gossip-based topology when a new node joins the cluster. When a boostrapping node starts gossiping, other nodes add it to their topology in `storage_service::on_alive`. Surprisingly, the state of the new node is set to `normal`, as it's the default value used by `add_or_update_endpoint`. Later, the state will be set to `bootstrapping` or `replacing`, and finally it will be set again to `normal` when the join operation finishes. We fix this strange behavior by setting the node state to `none` in `storage_service::on_alive` for nodes not present in the topology. Note that we must add such nodes to the topology. Other code needs their Host ID, IP, and location. We change the default node state from `normal` to `none` in `add_or_update_endpoint` to prevent bugs like the one in `storage_service::on_alive`. Also, we ensure that nodes in the `none` state are ignored in the getters of `locator::topology`.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	6adaf85634	test: boost: tablets tests: ensure all nodes are normal token owners In one of the following patches, we make NetworkTopologyStrategy and the tablet load balancer consider only normal token owners to ensure they ignore zero-token nodes. Some unit tests would start failing after this change because they do not ensure that all nodes are normal token owners. This patch prevents it. Judging by the logic in the test cases in `network_topology_strategy_test`, `point++` was probably intended anyway.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	366605224c	token_metadata: rename get_all_endpoints and get_all_ips In one of the following patches, we introduce support for zero-token nodes. A zero-token node that has successfully joined the cluster is in the normal state but is not a normal token owner. Hence, the names of `get_all_endpoints` and `get_all_ips` become misleading. They should specify that the functions return only IDs/IPs of token owners.	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	293a66fe41	network_topology_strategy: reallocate_tablets: remove unused dc_rack_nodes	2024-08-29 10:37:07 +02:00
Patryk Jędrzejczak	4ff08decb8	virtual_tables: cluster_status_table: execute: set dc regardless of the token ownership If a node is in `locator::topology`, then it has a location. We remove the token ownership condition to make the table more descriptive.	2024-08-29 10:37:06 +02:00
Kefu Chai	ecfe0aace6	perf: perf_mutation_readers: break memtable class down before this change, memtable serves as the fixture for 6 test cases, actually these 6 test cases can be categorized into a matrix of 3 x 2: { single_row, multi_row, large_partition } x { single_partition, multi_paritition }. in this change, we break memtable into 3 different fixtures, to reflect this fact. more readable this way. and a benefit is that each test does not have to pay for the overhead of setup it does not use at all. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20177	2024-08-29 08:54:17 +03:00
Botond Dénes	e538e3593c	Merge 'build: add --no-use-cmake option to configure.py' from Kefu Chai as part of the efforts to address scylladb/scylladb#2717, we are switching over to the CMake-based building system, and fade out the mechinary to create the rules manually in `configure.py`. in this change, we add `--no-use-cmake` to `configure.py`, it serves two purposes: * prepare for the change which enables cmake by default, by then, we would set the default value of `use_cmake` to True, and allow user to keep using the existing mechinary in the transition period using `--no-use-cmake`. * allows the CI to tell if a tree is able to build with CMake. the command line option of `--use-cmake` is also used by the CI workflows, and is passed to `configure.py` if `BUILD_WITH_CMAKE` jenkins pipeline parameter is set. but not all branches with `--use-cmake` are ready to build with CMake -- only the latest master HEAD is ready. so the CI needs to check the capability of building with CMake by looking at the output of `configure.py --help`, to see if it includes `--no-use-cmake`. after this change lands. we will remove the `BUILD_WITH_CMAKE` parameter, and use cmake as long as `configure.py` supports `--no-use-cmake` option. the existing mechinary will stay with us for a short transition period so that developers can take time to get used to the usage of the naming of targets and the new directory arrangement. as a side effect, #20079 will be fixed after switching to CMake. --- this is a cmake-related change, hence no need to backport. Closes scylladb/scylladb#20261 * github.com:scylladb/scylladb: build: add --no-use-cmake option to configure.py build: let configure.py fail if unknown option is passed to it	2024-08-29 08:51:41 +03:00
Kefu Chai	a182bfd96a	tools/read_mutation: reuse parse_table_directory_name() less repeatings this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20315	2024-08-29 08:49:20 +03:00
Nadav Har'El	6391550bbc	test/alternator: add another check to test_stream_list_tables The test test_streams.py::test_stream_list_tables reproduces a bug where enabling streams added a spurious result to ListTables. A reviewer of that patch asked to also add a check that name of the table itself doesn't disappear from ListTables when a stream is enabled, so this is what this patch adds. This theoretical scenario (a table's name disappearing from ListTables) never happened, so the new check doesn't reproduce any known bug, but I guess it never hurts to make the test stronger for regression testing. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#19934	2024-08-29 08:45:22 +03:00
Nadav Har'El	61e5927e8e	repair: fix build on older compilers The code tries to build as "neighbors" an unordered_map from an iterator of std::tuple, instead of the correct std::pair. Apparently, the tuples are transparently converted to pairs on the newest compilers and the whole works, but on slightly older compilers (like the one on Fedora 39) Scylla no longer compiles - the compiler complains it can't convert a tuple to a pair in this context. So fix the code to use pairs, not tuples, and it fixes the build on Fedora 39. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#20319	2024-08-28 19:56:03 +03:00
Laszlo Ersek	49bff3b1ab	generic_server: make server::stop() idempotent After server::shutdown(), make server::stop() more robust too, by allowing callers (internal or external) to call it several times (not concurrently though, just yet; see <https://github.com/scylladb/scylladb/issues/20309>). Suggested-by: Benny Halevy <bhalevy@scylladb.com> Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-28 15:54:31 +02:00
Kefu Chai	03ab80501f	tools/scylla-nodetool: add restore integration as we have an API for restore a keyspace / table, let's expose this feature with nodetool. so we can exercise it without the help of scylla-manager or 3rd-party tools with a user-friendly interface. in this change: * add a new subcommand named "restore" to nodetool * add test to verify its interaction with the API server * update the document accordingly. * the bash completion script is updated accordingly. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-28 15:42:49 +03:00
Pavel Emelyanov	41b9eda398	test/object_store: Add simple restore test The test shows how to restore previously backed up table: - backup - truncate to get rid of existing sstables - start restore with the new API method - wait for the task to finish Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-28 15:42:49 +03:00
Pavel Emelyanov	f5a22a94c6	test/object_store: Generalize prepare_snapshot_for_backup() Give it snapshot-name argument. Next test will want custom snapshot name. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-28 15:42:49 +03:00
Pavel Emelyanov	11a04bfb66	code: Introduce restore API method The method starts a task that uses sstables_loader load-and-stream functionality to bring new sstables into the cluster. The existing load-and-stream picks up sstables from upload/ directory, the newly introduced task collects them from S3 bucket and given prefix (that correspond to the path where backup API method put them). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-28 15:42:49 +03:00
Sergey Zolotukhin	65f37f3ba6	Ignore seed name resolution errors on restart. Gossiper seeds host name resolution failures are ignored during restart if a node is already boostrapped (i.e. it has successfully joined the cluster). Fixes scylladb/scylladb#14945	2024-08-28 14:01:04 +02:00
Patryk Jędrzejczak	08cb3a5e2c	test: test_raft_recovery_basic: add raft=trace logs It could help when we hit scylladb/scylladb#17918 again. This PR only changes log levels in a test, no need to backport it. Refs scylladb/scylladb#17918 Closes scylladb/scylladb#20318	2024-08-28 13:50:09 +02:00
Sergey Zolotukhin	fc5e683d02	Add a test for starting with a wrong seed. The test checks a bootstrapped node start with a wrong host name in the seeds config. Test for scylladb/scylladb#14945	2024-08-28 11:34:37 +02:00
Laszlo Ersek	1138347e7e	generic_server: coroutinize server::shutdown() By turning server::shutdown() into a coroutine, we need not dynamically allocate "nr_conn". Verified as follows: (1) In terminal #1: build/Dev/scylla --overprovisioned --developer-mode=yes \ --memory=2G --smp=1 --default-log-level error \ --logger-log-level cql_server=debug:cql_server_controller=debug > INFO [...] cql_server_controller - Starting listening for CQL clients > on 127.0.0.1:9042 (unencrypted, > non-shard-aware) > INFO [...] cql_server_controller - Starting listening for CQL clients > on 127.0.0.1:19042 (unencrypted, > shard-aware) (2) In terminals #2 and #3: tools/cqlsh/bin/cqlsh.py (3) Press ^C in terminal #1: > DEBUG [...] cql_server - abort accept nr_total=2 > DEBUG [...] cql_server - abort accept 1 out of 2 done > DEBUG [...] cql_server - abort accept 2 out of 2 done > DEBUG [...] cql_server - shutdown connection nr_total=4 > DEBUG [...] cql_server - shutdown connection 1 out of 4 done > DEBUG [...] cql_server - shutdown connection 2 out of 4 done > DEBUG [...] cql_server - shutdown connection 3 out of 4 done > DEBUG [...] cql_server - shutdown connection 4 out of 4 done > INFO [...] cql_server_controller - CQL server stopped This patch is best viewed with "git show --word-diff=color". Suggested-by: Benny Halevy <bhalevy@scylladb.com> Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-28 10:59:44 +02:00
Laszlo Ersek	2216275ebd	generic_server: make server::shutdown() idempotent Make server::shutdown() more robust by allowing callers (internal or external) to call it several times (not concurrently though, just yet; see <https://github.com/scylladb/scylladb/issues/20309>). Suggested-by: Benny Halevy <bhalevy@scylladb.com> Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-28 10:59:44 +02:00
Laszlo Ersek	dbc0ca6354	test/generic_server: add test case Check whether we can stop a generic server without first asking it to listen. The test fails currently; the failure mode is a hang, which triggers the 5 minute timeout set in the test: > unknown location(0): fatal error: in "stop_without_listening": > seastar::timed_out_error: timedout > seastar/src/testing/seastar_test.cc(43): last checkpoint > test/boost/generic_server_test.cc(34): Leaving test case > "stop_without_listening"; testing time: 300097447us Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-28 10:59:44 +02:00
Laszlo Ersek	931f2f8d73	configure, cmake: sort the lists of boost unit tests Both lists were obviously meant to be sorted originally, but by today we've introduced many instances of disorder -- thus, inserting a new test in the proper place leaves the developer scratching their head. Sort both lists. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-28 10:59:44 +02:00
Laszlo Ersek	5a04743663	generic_server: convert connection tracking to seastar::gate If we call server::stop() right after "server" construction, it hangs: With the server never listening (never accepting connections and never serving connections), nothing ever calls server::maybe_stop(). Consequently, co_await _all_connections_stopped.get_future(); at the end of server::stop() deadlocks. Such a server::stop() call does occur in controller::do_start_server() [transport/controller.cc], when - cserver->start() (sharded<cql_server>::start()) constructs a "server"-derived object, - start_listening_on_tcp_sockets() throws an exception before reaching listen_on_all_shards() (for example because it fails to set up client encryption -- certificate file is inaccessible etc.), - the "deferred_action" cserver->stop().get(); is invoked during cleanup. (The cserver->stop() call exposing the connection tracking problem dates back to commit `ae4d5a60ca` ("transport::controller: Shut down distributed object on startup exception", 2020-11-25), and it's been triggerable through the above code path since commit `6b178f9a4a` ("transport/controller: split configuring sockets into separate functions", 2024-02-05).) Tracking live connections and connection acceptances seems like a good fit for "seastar::gate", so rewrite the tracking with that. "seastar::gate" can be closed (and the returned future can be waited for) without anyone ever having entered the gate. NOTE: this change makes it quite clear that neither server::stop() nor server::shutdown() must be called multiple times. The permitted sequences are: - server::shutdown() + server::stop() - or just server::stop(). Fixes #10305 Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-28 10:59:44 +02:00
Kefu Chai	6d8dca1e20	build: add --no-use-cmake option to configure.py as part of the efforts to address scylladb/scylladb#2717, we are switching over to the CMake-based building system, and fade out the mechinary to create the rules manually in `configure.py`. in this change, we add `--no-use-cmake` to `configure.py`, it serves two purposes: * prepare for the change which enables cmake by default, by then, we would set the default value of `use_cmake` to True, and allow user to keep using the existing mechinary in the transition period using `--no-use-cmake`. * allows the CI to tell if a tree is able to build with CMake. the command line option of `--use-cmake` is also used by the CI workflows, and is passed to `configure.py` if `BUILD_WITH_CMAKE` jenkins pipeline parameter is set. but not all branches with `--use-cmake` are ready to build with CMake -- only the latest master HEAD is ready. so the CI needs to check the capability of building with CMake by looking at the output of `configure.py --help`, to see if it includes --no-use-cmake`. after this change lands. we will remove the `BUILD_WITH_CMAKE` parameter, and use cmake as long as `configure.py` supports `--no-use-cmake` option. the existing mechinary will stay with us for a short transition period so that developers can take time to get used to the usage of the naming of targets and the new directory arrangement. as a side effect, #20079 will be fixed after switching to CMake. Refs scylladb/scylladb#2717 Refs scylladb/scylladb#20079 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-28 11:37:56 +08:00
Kefu Chai	a2de14be7f	build: let configure.py fail if unknown option is passed to it this allows us to use `configure.py` to tell if a certain argument is supported without parsing its output. in the next commit, we will add `--no-use-cmake` option, which will be used to tell if the tree is ready for using CMake for its building system. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-28 11:37:55 +08:00
Kefu Chai	e4b213f041	build: cmake: use the same options to configure seastar in `configure.py`, a set of options are specified when configuring seastar, but not all of them were ported to scylla's CMake building system. for instance, `configure.py` explicitly disables io_uring reactor backend at build time, but the CMake-based system does not. so, in this change, in order to preserve the existing behavior, let's port the two previously missing option to CMake-based building system as well. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20288	2024-08-28 06:15:59 +03:00
Avi Kivity	94d5507237	Merge 'select from mutation_fragments() + tablets: handle reads for non-owned partitions' from Botond Dénes Attempting to read a partition via `SELECT * FROM MUTATION_FRAGMENTS()`, which the node doesn't own, from a table using tablets causes a crash. This is because when using tablets, the replica side simply doesn't handle requests for un-owned tokens and this triggers a crash. We should probably improve how this is handled (an exception is better than a crash), but this is outside the scope of this PR. This PR fixes this and also adds a reproducer test. Fixes: https://github.com/scylladb/scylladb/issues/18786 Fixes a regression introduced in 6.0, so needs backport to 6.0 and 6.1 Closes scylladb/scylladb#20109 * github.com:scylladb/scylladb: test/tablets: Test that reading tablets' mutations from MUTATION_FRAGMENTS works replica/mutation_dump: enfore pinning of effective replication map replica/mutation_dump: handle un-owned tokens (with tablets)	2024-08-27 20:46:10 +03:00
Avi Kivity	b13ab90448	Merge 'alternator/executor: Use native reversed format' from Łukasz Paszkowski When executing reversed queries, a native revered format shall be used. Therefore, the table schema and the clustering key bounds are reversed before a partition slice and a read command are constructed. It is, however, possible to run a reversed query passing a table schema but only when there are no restrictions on the clustering keys. In this particular situation, the query returns correct results. Since the current alternator tests in test.py do not imply any restrictions, this situation was not caught during development of https://github.com/scylladb/scylladb/pull/18864. Hence, additional tests are provided that add clustering keys restrictions when executing reversed queries to capture such errors earlier than in dtests. Additional manual tests were performed to test a mixed-node cluster (with alternator API enabled in Scylla on each node): 1. 2-node cluster with one node upgraded: reverse read queries performed on an old node 2. 2-node cluster with one node upgraded: reverse read queries performed on a new node 3. 2-node cluster with one node upgraded and all its sstable files deleted to trigger repair: reverse read queries performed on an old node 4. 2-node cluster with one node upgraded and all its sstable files deleted to trigger repair: reverse read queries performed on a new node All reverse read queries above consists of: - single-partition reverse reads with no clustering key restrictions, with single column restrictions and multi column restrictions both with and without paging turned on The exact same tests were also performed on a fully upgraded cluster. Fixes https://github.com/scylladb/scylladb/issues/20191 No backport is required as this is a complementary patch for the series https://github.com/scylladb/scylladb/pull/18864 that did not require backporting. Closes scylladb/scylladb#20205 * github.com:scylladb/scylladb: test_query.py: Test reverse queries with clustering key bounds alternator::do_query Add additional trace log alternator::do_query: Use native reversed format alternator::do_query Rename schema with table_schema	2024-08-27 20:40:49 +03:00
Benny Halevy	18c45f7502	raft_rebuild: propagate source_dc force option to rebuild_option Currently, the `force` property of the `source_dc` rebuild option is lost and `raft_topology_cmd_handler` has no way to know if it was given or not. This in turn can cause rebuild to fail, even when `--force` is set by the user, where it would succeed with gossip topology changes, based on the source_dc --force semantics. Fixes scylladb/scylladb#20242 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#20249	2024-08-27 17:05:48 +02:00
Kefu Chai	d27fdf9f57	Update seastar submodule * seastar a7d81328...83e6cdfd (29): > fair_queue: Export the number of times class was activated > tests/unit: drop support of C++17 > remove vestigial OSv support > cmake: undefine _FORTIFY_SOURCE on thread.cc > container_perf: a benchmark for container perf > io_sink: use chunked_fifo as _pending_io container > chunked_fifo: implement clear in terms of pop_n > chunked_fifo: pop_front_n > io_sink: use iteration instead of indexing > json2code_test: choose less popular port number > ioinfo: add '--max-reqsize' parameter > treewide: drop the support of fmtlib < 8.0.0 > build: bump up the required fmtlib version to 8.1.1 > conditional-variable: align when() and wait() behaviour in case of a predicate throwing an exception > stall-analyser: add output support for flamegraph > reactor: Add --io-completion-notify-ms option > io_queue: Stall detector > io_queue: Keep local variable with request execution delay > io_queue: Rename flow ratio timer to be more generic > reactor: Export _polls counter (internally) > dns: de-inline dns_resolver::impl methods > dns: enter seastar::net namespace > dnf: drop compatibility for c-ares <= 1.16 > reactor: add missing includes of noncopyable_function.hh > reactor: Reset one-shot signal to DFL before handling > future: correctly document nested exception type emitted by finally() > modules: fix FATAL_ERROR on compiler check > seastar.cc: include fmt/ranges.h > pack io_request Closes scylladb/scylladb#20300	2024-08-27 17:51:21 +03:00
Avi Kivity	2f4ef31254	Merge 'tools/testing: update dist-check to use rockylinux and adapt to cmake' from Kefu Chai `dist-check` tests the generated rpm packages by installing them in a centos 7 container. but this script is terribly outdated - centos 7 is deprecated. we should use a new distro's latest stable release. - cqlsh was added to the family of rpms a while ago. we should test it as well. - the directory hierarchy has been changed. we should read the artifacts from the new directories. - cmake uses a different directory hierarchy. we should check the directory used by cmake as well. to address these breaking changes, the scripts are updated accordingly. --- this change gives an overhaul to a test, which is not used in production. so no need to backport. Closes scylladb/scylladb#20267 * github.com:scylladb/scylladb: tools/testing: add cqlsh rpm tools/testing: adapt to cmake build directory tools/testing: test with rockylinux:9 not centos:7 tools/testing: correct the paths to rpm packages and SCYLLA-*-FILE dist-check: add :z option when mapping volume	2024-08-27 16:16:34 +03:00
Pavel Emelyanov	1f3f0b1926	sstable_loader: Add sstables::storage_manager dependency The storage_manager maintains set of clients to configured object storage(s). The sstables loader is going to spawn tasks that will talk to to those storages, thus it needs the storage manager to get the clients clients from. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:41 +03:00
Pavel Emelyanov	06c3c53deb	sstable_loader: Maintain task manager module This service is going to start tasks managed by task manager. For that, it should have its module set up and registered. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:41 +03:00
Pavel Emelyanov	9cf95e8a07	sstable_loader: Out-line constructor It will grow and become more complicated. Better to have it outside the header. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:41 +03:00
Pavel Emelyanov	6a006d2255	distributed_loader: Split get_sstables_from_upload_dir() Next patches will need this method to initialize sstable_directory differently and then do its regular processing. For that, split the method into two, next patch will re-use the common part it needs. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:41 +03:00
Pavel Emelyanov	630ab1dbea	sstables/storage: Compose uploaded sstable path simpler Current S3 storage driver keeps sstables in bucket in a form of /bucket/generation/component-name To get sstables that are backed up on S3 this format doesn't apply, because components are uploaded with their names unmodified. This patch makes S3 storage driver account for that and not re-format component paths for upload sstable state. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:41 +03:00
Pavel Emelyanov	2eda917375	sstable_directory: Prepare FS lister to scan files on S3 When component lister is created it checks the target storage options for what kind of lister to create. For local options it creates FS lister that collects sstables from their component files. For S3 options, it relies on sstables registry. When collecting sstables from backup, it's not possible to use registry, because those entries are not there. Instead, lister should pick up individual components as it they were on local FS. This patch prepares the lister for that -- in case S3 options are provided and the sstables' state is "upload", don't try to read those from registry, but instantiate the FS lister that will later use s3::bucket_lister. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:41 +03:00
Pavel Emelyanov	60d43911a9	sstable_directory: Parse sstable component without full path When sstable directory collects a entry from storage, it tries to parse its full path with the help of sstables::parse_path(). There are two overloads of that function -- one with ks:cf arguments and one without. The latter tries to "guess" keyspace and table names from the directory name. However, ks and table names are already known by the directory, it doesn't even use the returned ks and cf values, so this parsing is excessive. Also, future patches will put here backup paths, that might not match the ks_name/table_name-table_uuid/ pattern that the parser expects. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:41 +03:00
Pavel Emelyanov	86bc5b11fe	s3-client: Add support for lister::filter Directory lister comes with a filter function that tells lister which entries to skip by its .get() method. For uniformity, add the same to S3 bucket_lister. After this change the lister reports shorter name in the returned directory entry (with the prefix cut), so also need to tune up the unit test respectively. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:40 +03:00
Pavel Emelyanov	113d2449f8	utils: Introduce abstract (directory) lister This patch hides directory_lister and bucket_lister behind a common facade. The intention is to provide a uniform API for sstable_directory that it could use to list sstables' components wherever they are. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-27 16:15:40 +03:00
Piotr Dulikowski	da5f4faac1	Merge 'mv: reject user requests by coordinator when a replica is overloaded by MVs' from Wojciech Mitros Currently, when a view update backlog of one replica is full, the write is still sent by the coordinator to all replicas. Because of the backlog, the write fails on the replica, causing inconsistency that needs to be fixed by repair. To avoid these inconsistencies, this patch adds a check on the coordinator for overloaded replicas. As a result, a write may be rejected before being sent to any replicas and later retried by the user, when the replica is no longer overloaded. This patch does not remove the replica write failures, because we still may reach a full backlog when more view updates are generated after the coordinator check is performed and before the write reaches the replica. Fixes scylladb/scylladb#17426 Closes scylladb/scylladb#18334 * github.com:scylladb/scylladb: mv: test the view update behavior mv: add test for admission control storage_proxy: return overloaded_exception instead of throwing mv: reject user requests by coordinator when a replica is overloaded by MVs	2024-08-27 12:50:34 +02:00
Aleksandra Martyniuk	f38bb6483a	test: add test to ensure repair won't fail with uninitialized bm	2024-08-27 11:37:50 +02:00
Aleksandra Martyniuk	d8e4393418	repair: throw if batchlog manager isn't initialized repair_service::repair_flush_hints_batchlog_handler may access batchlog manager while it is uninitialized. Batchlog manager cannot be initialized before repair as we have the dependencies chain: repair_service -> storage_service::join_cluster -> batchlog_manager. Throw if batchlog manager isn't initialized. That won't cause repair to fail.	2024-08-27 11:22:28 +02:00
Botond Dénes	5c0f6d4613	Merge 'Make Summary support histogram with infinite bucket vlaues' from Amnon Heiman This series fixes an issue where histogram Summaries return an infinite value. It updated the quantile calculation logic to address cases where values fall into the infinite bucket of a histogram. Now, instead of returning infinite (max int), the calculation will return the last bucket limit, ensuring finite outputs in all cases. The series adds a test for summaries with a specific test case for this scenario. Fixes #20255 Need backport to 6.0, 6.1 and 2023.1 and above Closes scylladb/scylladb#20257 * github.com:scylladb/scylladb: test/estimated_histogram_test Add summary tests utils/histogram.hh: Make summary support inifinite bucket.	2024-08-27 10:33:54 +03:00
Kefu Chai	ae7ce38721	build: print out the default value of options instead of using the default `argparse.HelpFormatter`, let's use `ArgumentDefaultsHelpFormatter`, so that the default values of options are displayed in the help messages. this should help developer understand the behavior of the script better. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20262	2024-08-27 10:04:31 +03:00
Kefu Chai	e2747e4bb5	build: cmake: add dist-check target to achieve feature parity with our existing building system, we need to implement a new build target "dist-check" in the CMake-based building system. in this change, "dist-check" is added to CMake-based building system. unlike the rules generated by `configure.py`, the `dist-check` target in CMake depends on the dist-*-rpm targets. the goal is to enable user to test `dist-check` without explicitly building the artifacts being tested. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20266	2024-08-27 10:03:41 +03:00
Kefu Chai	ea612e7065	docs: install poetry>=1.8.0 in `57def6f1`, we specified "package-mode" for poetry, but this option was introduced in poetry 1.8.0, as the "non-package" mode support. see https://github.com/python-poetry/poetry/releases/tag/1.8.0 this change practically bumps up the minimum required poetry version to 1.8.0, we did update `pyproject.tombl` to reflect this change. but wefailed to update the `Makefile`. in this change, we update `Makefile` to ensure that user which happens have an older version of poetry can install the version which supports this version when running `make setupenv`. Refs scylladb/scylladb#20284 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20286	2024-08-27 09:20:09 +03:00
Yaniv Michael Kaul	022eb25d98	tools/toolchain/README.md: fix wording Forgot to add that 'reg' tool is also needed. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Closes scylladb/scylladb#20287	2024-08-27 09:18:23 +03:00
Kefu Chai	5cffb23aa3	scylla-gdb.py: use chunked_fifo to represent _sink._pending_io we switched from `circular_buffer` to `chunked_fifo` to present `io_sink::_pending_io` in the latest seastar now. to be prepared for this change, let's * add `chunked_fifo` class in `scylla-gdb.py`. * use `circular_buffer` as a fallback of `chunked_fifo`. instead of doing this the other way around, we try to send the message that the latest seastar uses `chunked_fifo`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20280	2024-08-27 08:44:56 +03:00
Andrei Chekun	fd51332978	test.py: Add parameter to control the pool size from the command line Add parameter --cluster-pool-size that can control pool size for all PythonTestSuite tests. By default, the pool size set to 10 for most of the suites, but this is too much for laptops. So this parameter can be used to lower the pool size and not to freeze the system. Additionally, the environment variable CLUSTER_POOL_SIZE was added for a convenient way to limit pool size in the system without the need to provide each time an additional parameter. Related: https://github.com/scylladb/scylladb/pull/20276 Closes scylladb/scylladb#20289	2024-08-26 19:55:41 +03:00
Avi Kivity	0acfa4a00d	Merge 'abstract_replication_strategy: make get_ranges async' from Benny Halevy To prevent stalls due to large number of tokens. For example, large cluster with say 70 nodes can have more than 16K tokens. Fixes #19757 Closes scylladb/scylladb#19758 * github.com:scylladb/scylladb: abstract_replication_strategy: make get_ranges async database: get_keyspace_local_ranges: get vnode_effective_replication_map_ptr param compaction: task_manager_module: open code maybe_get_keyspace_local_ranges alternator: ttl: token_ranges_owned_by_this_shard: let caller make the ranges_holder alternator: ttl: can pass const gms::gossiper& to ranges_holder alternator: ttl: ranges_holder_primary: unconstify _token_ranges member alternator: ttl: refactor token_ranges_owned_by_this_shard	2024-08-26 16:56:18 +03:00
Botond Dénes	6d633e89ef	Merge 'update CODEOWNERS' from Piotr Smaron Removed people that no longer contribute to the scylladb.git and added/substituted reviewers responsible for maintaining the frontend components. No need to backport, this is just an information for the github tool. Closes scylladb/scylladb#20136 * github.com:scylladb/scylladb: codeowners: add appropriate reviewers to the cluster components codeowners: add appropriate reviewers to the frontend components codeowners: fix codeowner names codeowners: remove non contributors	2024-08-26 16:44:39 +03:00
Botond Dénes	4505b14fd6	Merge 'table_helper: complete coroutinization' from Avi Kivity table_helper has some quite awkward code, improve it a little. Code cleanup, so no reason to backport. Closes scylladb/scylladb#20194 * github.com:scylladb/scylladb: table_helper: insert(): improve indentation table_helper: coroutinize insert() table_helper: coroutinize cache_table_info() table_helper: extract try_prepare()	2024-08-26 13:43:17 +03:00
Botond Dénes	b2c07c9b6f	Merge 'compaction: change compaction stop reason ' from Aleksandra Martyniuk Currently "table removal" is logged as a reason of compaction stop for table drop, tablet cleanup and tablet split. Modify log to reflect the reason. Closes scylladb/scylladb#20042 * github.com:scylladb/scylladb: test: add test to check compaction stop log compaction: fix compaction group stop reason	2024-08-26 13:40:07 +03:00
Kefu Chai	4d516a8363	tools/testing: add cqlsh rpm we need to test the installation of cqlsh rpm. also, we should use the correct paths of the generated rpm packages. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-26 11:33:57 +08:00
Kefu Chai	baee15390e	tools/testing: adapt to cmake build directory cmake uses a different arrangement, so let's check for the existence of the build directory and fallback to cmake's build directory. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-26 11:33:57 +08:00
Kefu Chai	b802c000e1	tools/testing: test with rockylinux:9 not centos:7 the centos image repos on docker has been deprecated, and the repo for centos7 has been removed from the main CentOS servers. so we are either not able to install packages from its default repo, without using the vault mirror, or no longer to pull its image from dockerhub. so, in this change * we switch over to rockylinux:9, which is the latest stable release of rockylinux, and rockylinux is a popular clone of RHEL, so it matches our expectation of a typical use case of scylla. * use dnf to manage the packages. as dnf is the standard way to manage rpm packages in modern RPM-based distributions. * do not install deltarpm. delta rpms are was not supported since RHEL8, and the `deltarpm` package is not longer available ever since. see https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/8/html-single/considerations_in_adopting_rhel_8/index#ref_the-deltarpm-functionality-is-no-longer-supported_notable-changes-to-the-yum-stack as a sequence, this package does not exist in Rockylinux-9. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-26 11:33:53 +08:00
Kefu Chai	00dad27f67	tools/testing: correct the paths to rpm packages and SCYLLA-*-FILE when building with the rules generated from `configure.py`, these files are located under tools' own build directory. so correct them. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-26 11:19:24 +08:00
Kefu Chai	86ef63df92	dist-check: add :z option when mapping volume if SELinux is enabled on the host, we'd have following failure when running `dist-check.sh`: ``` + podman run -i --rm -v /home/kefu/dev/scylladb:/home/kefu/dev/scylladb docker.io/centos:7 /bin/bash -c 'cd /home/kefu/dev/scylladb && /home/kefu/dev/scylladb/tools/testing/dist-check/docker.io/centos-7.sh --mode debug' /bin/bash: line 0: cd: /home/kefu/dev/scylladb: Permission denied ``` to address the permission issue, we need to instruct podman to relabel the shared volume, so that the container can access the shared volume. see also https://docs.podman.io/en/stable/markdown/podman-pod-create.1.html#volume-v-source-volume-host-dir-container-dir-options Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-26 11:15:40 +08:00
Kefu Chai	8ef26a9c8c	build: cmake: add "test" target before this change, none of the target generated by CMake-based building system runs `test.py`. but `build.ninja` generated directly by `configure.py` provides a target named `test`, which runs the `test.py` with the options passed to `configure.py`. to be more compatible with the rules generated by `configure.py`, in this change * do not include "CTest" module, as we are not using CTest for driving tests. we use the homebrew `test.py` for this purpose. more importantly, the target named "test" is provided by "CTest". so in order to add our own "test" target, we cannot use "CTest" module. * add a target named "test" to run "test.py". * add two CMake options so we can customize the behavior of "test.py", this is to be compatible with the existing behavior of `configure.py`. Refs scylladb/scylladb#2717 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20263	2024-08-25 21:45:13 +03:00
Avi Kivity	72a85e3812	Merge 'Integrated backup' from Pavel Emelyanov This adds minimal implementation of the start-backup API call. The method starts a task that uploads all files from the given keyspace's snapshot to the requested endpoint/bucket. Arguments are: - endpoint -- the ID in object_store.yaml config file - bucket -- the target bucket to put objects into - keyspace -- the keyspace to work on - snapshot -- the method assumes that the snapshot had been already taken and only copies sstables from it The task runs in the background, its task_id is returned from the method once it's spawned and it should be used via /task_manager API to track the task execution and completion (hint: it's good to have non-zero TTL value to make sure fast backups don't finish before the caller manages to call wait_task API). Sstables components are scanned for all tables in the keyspace and are uploaded into the /bucket/${cf_name}/${snapshot_name}/ path. refs: #18391 Closes scylladb/scylladb#19890 * github.com:scylladb/scylladb: tools/scylla-nodetool: add backup integration docs: Document the new backup method test/object_store: Test that backup task is abortable test/object_store: Add simple backup test test/object_store: Move format_tuples() test/pylib: Add more methods to rest client backup-task: Make it abortable (almost) code: Introduce backup API method database: Export parse_table_directory_name() helper database: Introduce format_table_directory_name() helper snapshot-ctl: Add config to snapshot_ctl snapshot-ctl: Add sstables::storage_manager dependency snapshot-ctl: Maintain task manager module snapshot-ctl: Add "snapshots" logger snapshot-ctl: Outline stop() method and constructor snapshot-ctl: Inline run_snapshot_list<> test/cql_test_env: Export task manager from cql test env task_manager: Print task ttl on start (for debugging) docs: Update object_storage.md with AWS_ environment docs: Restructure object_storage.md	2024-08-25 20:19:10 +03:00
Kefu Chai	f8931a4578	build: cmake: add "dist" target since the rules generated by `configure.py` has this target, we need to have an equivalent target as well in CMake-based buidling system. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20265	2024-08-25 20:18:12 +03:00
Andrei Chekun	f54b7f5427	test.py: Increase pool size Increase pool size changes were recently reverted because of the flakiness for the test_gossip_boot test. Test started to fail on adding the node to the cluster without any issues in the Scylla log file. In test logs it looked like the installation process for the new node just hanged. After investigating the problem, I've found out that the issue is that test.py was draining the io_executor pool for cleaning the directory during install that was set to eight workers. So to fix the issue, io_executor pool should be increased to more or less the same ratio as it was: doubled cluster pool size. Closes scylladb/scylladb#20276	2024-08-25 19:59:18 +03:00
Kefu Chai	a0688b29ea	replication_strategy: add fmt::formatter<replication_strategy_type> so that we can use {fmt} with it without the help of fmt::streamed. also since we have a proper formatter for replication_strategy_type, let's implement `formatter<vnode_effective_replication_map::factory_key>` as well. since there are no callers of these two operator<<, let's drop them in this change. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20248	2024-08-25 19:34:52 +03:00
Kefu Chai	c88b63ce13	github: use clang-20 in clang-nightly workflow since clang 19 has been branched. let's track the development brach, which is clang 20. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20279	2024-08-25 19:31:43 +03:00
Benny Halevy	686a8f2939	abstract_replication_strategy: make get_ranges async To prevent stalls due to large number of tokens. For example, large cluster with say 70 nodes can have more than 16K tokens. Fixes #19757 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-25 10:57:34 +03:00
Benny Halevy	2bbbe2a8bc	database: get_keyspace_local_ranges: get vnode_effective_replication_map_ptr param Prepare for making the function async. Then, it will need to hold on to the erm while getting the token_ranges asynchronously. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-25 10:55:33 +03:00
Benny Halevy	ea5a0cca10	compaction: task_manager_module: open code maybe_get_keyspace_local_ranges It is used only here and can be simplified by checking if the keyspace replication strategy is per table by the caller. Prepare for making get_keyspace_local_ranges async. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-25 10:25:32 +03:00
Benny Halevy	824bdf99d2	alternator: ttl: token_ranges_owned_by_this_shard: let caller make the ranges_holder Add static `make` methods to ranges_holder_{primary,secondary} and use them to make the ranges objects and pass them to `token_ranges_owned_by_this_shard`, rather than letting token_ranges_owned_by_this_shard invoke the right constructor of the ranges_holder class. Prepare for making `make` async. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-25 10:25:32 +03:00
Benny Halevy	b2abbae24b	alternator: ttl: can pass const gms::gossiper& to ranges_holder There's no need to pass a mutable reference to the gossiper. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-25 10:25:32 +03:00
Benny Halevy	333c0d7c88	alternator: ttl: ranges_holder_primary: unconstify _token_ranges member To allow the class to be nothrow_move_constructable. Prepare for returning it as a future value. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-25 10:25:32 +03:00
Benny Halevy	d385219a12	alternator: ttl: refactor token_ranges_owned_by_this_shard Rather than holding a variant member (and defining both ranges_holder_{primary,secondary} in both specilizations of the class, just make the internal ranges_holder class first-class citizens and parameterize the `token_ranges_owned_by_this_shard` template by this class type. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-25 10:25:32 +03:00
Avi Kivity	c4dd21de38	repair: row_level: coroutinize repair_reader::close()	2024-08-24 00:36:48 +03:00
Avi Kivity	b1dd470533	repair: row_level: coroutinize repair_reader::end_of_stream()	2024-08-24 00:35:59 +03:00
Avi Kivity	7ce76fd0ea	repair: row_level: coroutinize sink_source_for_repair::close() The repeat() loop translates to almost nothing.	2024-08-24 00:30:02 +03:00
Avi Kivity	168a018e45	repair: row_level: coroutinize sink_source_for_repair::get_sink_source()	2024-08-24 00:19:12 +03:00
Avi Kivity	6b370d8154	table_helper: insert(): improve indentation Restore after coroutinization.	2024-08-24 00:08:05 +03:00
Avi Kivity	ecd7702007	table_helper: coroutinize insert() Improves readability. The do_with() ensures it's at least as performant (though it's not in any fast path).	2024-08-24 00:08:05 +03:00
Avi Kivity	980ec2f925	table_helper: coroutinize cache_table_info() After we extracted try_prepare(), this is fairly simple, and improves readability.	2024-08-24 00:08:05 +03:00
Avi Kivity	4e44a15d4d	table_helper: extract try_prepare() table_helper::cache_table_info() is fairly convoluted. It cannot be easily coroutinized since it invokes asynchronous functions in a catch block, which isn't supported in coroutines. To start to break it down, extract a block try_prepare() from code that is called twice. It's both a simplification and a first step towards coroutinization. The new try_prepare() can return three values: `true` if it succeeded, `false` if it failed and there's the possibility of attempting a fallback, and an exception on error.	2024-08-24 00:08:05 +03:00
Lakshmi Narayanan Sreethar	4823a1e203	test/pylib: fix keyspace_compaction method The `keyspace_compaction` method incorrectly appends the column family parameter to the URL using a regular string, `"?cf={table}"`, instead of an f-string, `f"?cf={table}"`. As a result, the column family name is sent as `{table}` to the server, causing the compaction request to fail. Fix this issue by passing the parameter to the POST request using a dictionary instead of appending it to the URL. Fixes #20264 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com> Closes scylladb/scylladb#20243	2024-08-23 15:20:10 +03:00
Kefu Chai	4a405b0af9	perf/perf_sstable: enumerate sstables when loading them before this change, we use the default options when creating `test_env`, and the default options enable `use_uuid`. but the modes of `perf-sstables` involving reads assumes that the identifiers are deterministic. so that the previously written sstables using the "write" mode can be read with the modes like "index_read", which just uses `test_env::make_sstable()` in `load_sstables()`, and under the hood, `test_env::make_sstable()` uses `test_env::new_generation()` for retrieving the next identifier of sstable. when using integer-base identifier, this works. as the sstable identifiers are generated from a monotonically increasing integer sequence, where the identifiers are deterministic. but this does not apply anymore when the UUID-based identifiers are used, as the identifiers are generated with a pseudorandom generator of UUID v1. in this change, to avoid relying on the determinism of the integer-based sstable identifier generation, we enumerate sstables by listing the given directory, and parse the path for their identifier. after this change, we are able to support the UUID-based sstable identifier. another option is disable the UUID-based sstable identifier when loading sstables. the upside is that this approach is minimal and straightforward. but the downside is that it encodes the assumption in the algorithm implicitly, and could be confusing -- we create a new generation for loading an existing sstable with this generation. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20183	2024-08-23 10:39:24 +03:00
Pavel Emelyanov	d1ac58f088	api: Get compaction througput via compaction manager Now the endpoint hanler gets the value from db::config which is not nice from several perspectives. First, it gets config (ab)using database. Second, it's compaction manager that "knows" its throughput, global config is the initial source of that information. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20173	2024-08-23 10:33:03 +03:00
Pavel Emelyanov	38edbebb10	compaction_manager: Keep flush-all-before-major option on own config Currently the major compaction task impl grabs this (non-updateable) value from db::config. That's not good, all services including compaction manager have their own configs from which they take options. Said that, this patch puts the said option onto compaction_manager::config, makes use of it and configures one from db::config on start (and tests). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20174	2024-08-23 10:31:55 +03:00
Botond Dénes	15fdc3f6cc	Merge 'Add ability to list S3 bucket contents' from Pavel Emelyanov This is prerequisite for "restore from object storage" feature. In order to collect the sstables in bucket one would need to list the bucket contents with the given prefix. The ListObjectsV2 provides a way for it and here's the respective s3::client extension. Closes scylladb/scylladb#20120 * github.com:scylladb/scylladb: test: Add test for s3::client::bucket_lister s3_client: Add bucket lister s3_client: Encode query parameter value for query-string	2024-08-23 10:16:07 +03:00
Kefu Chai	7f65ee3270	dbuild: pass --tty only if --interactive in `947e2814`, we pass `--tty` as long as we are using podman _or_ we are in interactive mode. but if we build the tree using podman using jenkins, we are seeing that ninja is displaying the output as if it's in an interactive mode. and the output includes ASCII escape codes. this is distracting. the reason is that we * are using podman, and * ninja tells if it should displaying with a "smart" terminal by checking istty() and the "TERM" environmental variable. so, in this change, we add --tty only if * we are in the interactive mode. * or stdin is associated with a terminal. this is the use case where user uses dbuild to interactively build scylla Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20196	2024-08-23 09:30:20 +03:00
Kefu Chai	ee19bbed05	test: do not define boost_test_print_type() for types with operator<< in `30e82a81`, we add a contraint to the template parameter of boost_test_print_type() to prevent it from being matched with types which can be formatted with operator<<. but it failed to work. we still have test failure reports like: ``` [Exception] - critical check ['s', 's', 't', '_', 'm', 'r', '.', 'i', 's', '_', 'e', 'n', 'd', '_', 'o', 'f', '_', 's', 't', 'r', 'e', 'a', 'm', '(', ')'] has failed ``` this is not what we expect. the reason is that we passed the template parameters to the `has_left_shift` trait in the wrong order, see https://live.boost.org/doc/libs/1_83_0/libs/type_traits/doc/html/boost_typetraits/reference/has_left_shift.html. we should have passed the lhs of operator<< expression as first parameter, and rhs the second. so, in this change, we correct the type constraint by passing the template parameter in the right order, now the error message looks better, like: ``` test/boost/mutation_query_test.cc(110): error: in "test_partition_query_is_full": check !partition_slice_builder(*s) .with_range({}) .build() .is_full() has failed ``` it turns out boost::transformed_range<> is formattable with operator<<, as it fulfills the constraints of `boost::has_left_shift<ostream, R>`, but when printing it, the compiler fails when it tries to insert the elements in the range to the output stream. so, in order to workaround this issue, we add a specialization for `boost::transformed_range<F, R`. also, to improve the readability, we reimplement the `has_left_shift<>` as a concept, so that it's obvious that we need to put both the output stream as the first parameter. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20233	2024-08-23 09:26:22 +03:00
Amnon Heiman	644e6f0121	test/estimated_histogram_test Add summary tests This patch adds tests for summary calculation. It adds two tests, the first is a basic calculation for P50, P95, P99 by adding 100 elements into 20 buckets. The second test look that if elements are found in the infinite bucket, the result would be the lower limit (33s) and not infinite. Relates to #20255 Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-08-22 23:34:24 +03:00
Amnon Heiman	011aa91a8c	utils/histogram.hh: Make summary support inifinite bucket. This patch handles an edge cases related to The infinite bucket limit. Summaries are the P50, P95, and P99 quantiles. The quantiles are calculated from a histogram; we find the bucket and return its upper limit. In classic histograms, there is a notion of the infinite bucket; anything that does not fall into the last bucket is considered to be infinite; with quantile, it does not make sense. So instead of reporting infinite we'll report the bucket lower limit. Fixes #20255 Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-08-22 23:34:24 +03:00
Kefu Chai	39dd088374	test: include used headers before this change, clang 20 fails to build the tree, like: ``` /home/kefu/.local/bin/clang++ -DBOOST_ALL_DYN_LINK -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TESTING_MAIN -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -isystem /home/kefu/dev/scylladb/abseil -isystem /home/kefu/dev/scylladb/build/rust -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -U_FORTIFY_SOURCE -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -fno-sanitize=vptr -MD -MT test/boost/CMakeFiles/database_test.dir/Debug/database_test.cc.o -MF test/boost/CMakeFiles/database_test.dir/Debug/database_test.cc.o.d -o test/boost/CMakeFiles/database_test.dir/Debug/database_test.cc.o -c /home/kefu/dev/scylladb/test/boost/database_test.cc /home/kefu/dev/scylladb/test/boost/database_test.cc:539:29: error: invalid use of incomplete type 'schema_builder' 539 \| return *schema_builder(ks_name, cf_name) \| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ /home/kefu/dev/scylladb/schema/schema.hh:115:7: note: forward declaration of 'schema_builder' 115 \| class schema_builder; \| ^ ``` and ``` /home/kefu/.local/bin/clang++ -DBOOST_ALL_DYN_LINK -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TESTING_MAIN -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -isystem /home/kefu/dev/scylladb/abseil -isystem /home/kefu/dev/scylladb/build/rust -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -U_FORTIFY_SOURCE -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -fno-sanitize=vptr -MD -MT test/boost/CMakeFiles/group0_cmd_merge_test.dir/Debug/group0_cmd_merge_test.cc.o -MF test/boost/CMakeFiles/group0_cmd_merge_test.dir/Debug/group0_cmd_merge_test.cc.o.d -o test/boost/CMakeFiles/group0_cmd_merge_test.dir/Debug/group0_cmd_merge_test.cc.o -c /home/kefu/dev/scylladb/test/boost/group0_cmd_merge_test.cc /home/kefu/dev/scylladb/test/boost/group0_cmd_merge_test.cc:78:18: error: member access into incomplete type 'db::config' 78 \| cfg.db_config->commitlog_segment_size_in_mb(1); \| ^ /home/kefu/dev/scylladb/data_dictionary/data_dictionary.hh:28:7: note: forward declaration of 'db::config' 28 \| class config; \| ^ 1 error generated. ``` and ``` `FAILED: test/boost/CMakeFiles/repair_test.dir/Debug/repair_test.cc.o /home/kefu/.local/bin/clang++ -DBOOST_ALL_DYN_LINK -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TESTING_MAIN -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/build/gen -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -isystem /home/kefu/dev/scylladb/abseil -isystem /home/kefu/dev/scylladb/build/rust -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -U_FORTIFY_SOURCE -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -fno-sanitize=vptr -MD -MT test/boost/CMakeFiles/repair_test.dir/Debug/repair_test.cc.o -MF test/boost/CMakeFiles/repair_test.dir/Debug/repair_test.cc.o.d -o test/boost/CMakeFiles/repair_test.dir/Debug/repair_test.cc.o -c /home/kefu/dev/scylladb/test/boost/repair_test.cc /home/kefu/dev/scylladb/test/boost/repair_test.cc:149:45: error: use of undeclared identifier 'global_schema_ptr' 149 \| co_await e.db().invoke_on_all([gs = global_schema_ptr(gen.schema())](replica::database& db) -> future<> { \| ^ /home/kefu/dev/scylladb/test/boost/repair_test.cc:150:62: error: use of undeclared identifier 'gs' 150 \| co_await db.add_column_family_and_make_directory(gs.get(), replica::database::is_new_cf::yes); \| ^ 2 errors generated. ``` because we are using incomplete types when their complete definitions are required. so, in this change, we include the headers for their complete definition. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20239	2024-08-22 20:51:38 +03:00
Kefu Chai	969cbb75ce	tools/scylla-nodetool: add backup integration as we have an API for backup a keyspace, let's expose this feature with nodetool. so we can exercise it without the help of scylla-manager or 3rd-party tools with a user-friendly interface. in this change: * add a new subcommand named "backup" to nodetool * add test to verify its interaction with the API server * add two more route to the REST API mock server, as the test is using /task_manager/wait_task/{task_id} API. for the sake of completeness, the route for /task_manager/{part1} is added as well. * update the document accordingly. * the bash completion script is updated accordingly. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-22 19:48:06 +03:00
Pavel Emelyanov	245cc852dd	docs: Document the new backup method Add the new /storage_service/backup endpoint to object_storage.md as yet another way to use S3 from Scylla.	2024-08-22 19:47:06 +03:00
Pavel Emelyanov	de87450453	test/object_store: Test that backup task is abortable It starts similarly to simpl backup test, but injects a pause into the task once a single file is scheduled for upload, then aborts the task, waits for it to fail, and check that _not_ all files are uploaded. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 19:47:06 +03:00
Pavel Emelyanov	f8d894bc23	test/object_store: Add simple backup test The test shows how to backup a keyspace: - flush - take snapshot - start backup with the new API method - wait for the task to finish Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 19:47:06 +03:00
Pavel Emelyanov	47e49e6dec	test/object_store: Move format_tuples() There will soon appear a new .py file in the suite that will want to use this helper too Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 19:47:06 +03:00
Pavel Emelyanov	d83d585709	test/pylib: Add more methods to rest client Namely: - POST /storage_service/snapshots to take snapshot on a ks - GET /task_manager/get_task_status/{id} to get status of a running task - GET /task_manager/wait_task/{id} to wait for a task to finish - POST /task_manager/abort_task/{id} to abort a running task Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 19:47:06 +03:00
Pavel Emelyanov	ed6e6700ab	backup-task: Make it abortable (almost) Make the impl::is_abortable() return 'yes' and check the impl::_as in the files listing loop. It's not real abort, since files listing loop is expected to be fast and most of the time will be spent in s3::client code reading data from disk and sending them to S3, but client doesn't support aborting its requests. That's some work yet to be done. Also add injection for future testing. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 19:47:06 +03:00
Pavel Emelyanov	a812f13ddd	code: Introduce backup API method The method starts a task that uploads all files from the given keyspace's snapshot to the requested endpoint/bucket. The task runs in the background, its task_id is returned from the method once it's spawned and it should be used via /task_manager API to track the task execution and completion (hint: it's good to have non-zero TTL value to make sure fast backups don't finish before the caller manages to call wait_task API). If snapshot doesn't exist, nothing happens (FIXME, need to return back an error in that case). If endpoint is not configured locally, the API call resolves with bad-request instantly. Sstables components are scanned for all tables in the keyspace and are uploaded into the /bucket/${cf_name}/${snapshot_name}/ path. Task is not abortable (FIXME -- to be added) and doesn't really report its progress other than running/done state (FIXME -- to be added too). Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 19:47:06 +03:00
Pavel Emelyanov	f7b380d53b	database: Export parse_table_directory_name() helper There's parse_table_directory_name() static helper in database.cc code that is used by methods that parse table tree layout for snapshot. Export this helper for external usage and rename to fit the format_... one introduced by previous patch. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:57:48 +03:00
Pavel Emelyanov	33962946fc	database: Introduce format_table_directory_name() helper The one makes table directory (not full path) out of table name and uuid. This is to be symmetrical with yet another helper that converts dirctory name back to table name and uuid (next patch) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:57:48 +03:00
Pavel Emelyanov	dff51fd58c	snapshot-ctl: Add config to snapshot_ctl Pretty much all services in Scylla have their own config. Add one to snapshot-ctl too, it will be populated later. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:57:20 +03:00
Pavel Emelyanov	f37857e20a	snapshot-ctl: Add sstables::storage_manager dependency The storage_manager maintains set of clients to configured object storage(s). The snapshot ctl is going to spawn tasks that will talk to those storages, thus it needs the storage manager to get the clients from. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	362331c89b	snapshot-ctl: Maintain task manager module This service is going to start tasks managed by task manager. For that, it should have its module set up and registered. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	4ae89a9c81	snapshot-ctl: Add "snapshots" logger Will be used later Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	90c794172b	snapshot-ctl: Outline stop() method and constructor These two are going to grow, keep them out not to pollute the header Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	96946a4b11	snapshot-ctl: Inline run_snapshot_list<> This helper will be used by a code from another .cc file, so the template needs to be in header for smooth instantiation Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	4e73b4d8ad	test/cql_test_env: Export task manager from cql test env To be used by one of the next patches Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	4b86eede1f	task_manager: Print task ttl on start (for debugging) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	8949d73cd9	docs: Update object_storage.md with AWS_ environment Commit `51c53d8db6` made it possible to configure object storage endpoint creds via environment. Mention this in the docs.	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	d3f9865d2f	docs: Restructure object_storage.md Currently the doc assumes that object storage can only be used to keep sstables on it. It's going to change, restructure the doc to allow for more usage scenarios.	2024-08-22 14:08:21 +03:00
Pavel Emelyanov	4e2d7aa2a2	test/tablets: Test that reading tablets' mutations from MUTATION_FRAGMENTS works Currently it doesn't, one of the node crashes with std::out_of_range exception and meaningless calltrace [Botond]: this test checks the case of reading a partition via MUTATION_FRAGMENTS from a node which doesn't own said partition. refs: #18786 Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-22 06:24:06 -04:00
Botond Dénes	46563d719f	replica/mutation_dump: enfore pinning of effective replication map By making it a required argument, making sure the topology version is pinned for the duration of the query. This is needed because mutation dump queries bypass the storage proxy, where this pinning usually takes place. So it has to be enforced here.	2024-08-22 06:24:06 -04:00
Botond Dénes	de5329157c	replica/mutation_dump: handle un-owned tokens (with tablets) When using tablets, the replica-side doesn't handle un-owned tokens. table::shard_for_reads() will just return 0 for un-owned tokens, and a later attempt at calling table::storage_group_for_token() with said un-owned token will cause a crash (std::terminate due to std::out_of_range thrown in noexcept context). The replicas rely on the coordinator to not send stray requests, but for select from mutation_fragments(table) queries, there is no coordinator side who could do the correct dispatching. So do this in mutation_dump(), just creating empty readers for un-owned tokens.	2024-08-22 03:06:55 -04:00
Łukasz Paszkowski	a11d19f321	test_query.py: Test reverse queries with clustering key bounds Since a native reversed format is used for reversed queries, additional tests with restrictions on clustering keys are required to capture possible errors like https://github.com/scylladb/scylladb/issues/20191 earlier than in dtests. Add parametrization to the following tests: + test_query_reverse + test_query_reverse_paging to accept a comparison operator used in selection criteria for a Query operation.	2024-08-21 14:21:34 +02:00
Aleksandra Martyniuk	9b7c837106	test: add test to check compaction stop log	2024-08-21 12:42:37 +02:00
Aleksandra Martyniuk	5005e19de7	compaction: fix compaction group stop reason compaction_manager::remove passes "table removal" as a reason of stopping ongoing compactions, but currently remove method is also called when a tablet is migrated or split. Pass the actual reason of compaction stop, so that logs aren't misleading.	2024-08-21 12:42:09 +02:00
Avi Kivity	2ef5b5e4fe	Revert "[test.py] Increase pool size for CI" This reverts commit `cc428e8a36`. It causes may spurious CI failures while nodes are being torn down. Revert it until the root cause is fixed, after which it can be reinstated. Fixes #20116.	2024-08-21 13:21:08 +03:00
Benny Halevy	f40d06b766	table: calculate_tablet_count: use sg_manager storage_groups size Now, when each shard storage_group_manager keeps only the storage_groups for the tablet replica it owns, we can simple return the storage_group map size instead of counting the number of tablet replicas mapped to this shard. Add a unit test that sums the tablet count on all shards and tests that the sum is equal to the configured default `initial_tablets. Fixes #18909 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#20223	2024-08-21 11:01:58 +02:00
Tomasz Grabiec	a3a97e8aad	Merge 'schema_tables: calculate_schema_digest: prevent stalls due to large m…' from Benny Halevy …utations vector With a large number of table the schema mutations vector might get big enoug to cause reactor stalls when freed. For example, the following stall was hit on 2023.1.0~rc1-20230208.fe3cc281ec73 with 5000 tables: ``` (inlined by) ~vector at /usr/bin/../lib/gcc/x86_64-redhat-linux/12/../../../../include/c++/12/bits/stl_vector.h:730 (inlined by) db::schema_tables::calculate_schema_digest(seastar::sharded<service::storage_proxy>&, enum_set<super_enum<db::schema_feature, (db::schema_feature)0, (db::schema_feature)1, (db::schema_feature)2, (db::schema_feature)3, (db::schema_feature)4, (db::schema_feature)5, (db::schema_feature)6, (db::schema_feature)7> >, seastar::noncopyable_function<bool (std::basic_string_view<char, std::char_traits<char> >)>) at ./db/schema_tables.cc:799 ``` This change returns a mutations generator from the `map` lambda coroutine so we can process them one at a time, destroy the mutations one at a time, and by that, reducing memory footprint and preventing reactor stalls. Fixes #18173 Closes scylladb/scylladb#18174 * github.com:scylladb/scylladb: schema_tables: calculate_schema_digest: filter the key earlier schema_tables: calculate_schema_digest: prevent stalls due to large mutations vector	2024-08-20 21:24:38 +02:00
Łukasz Paszkowski	f29d7ffa81	alternator::do_query Add additional trace log Additional log prints information on the read query being executed. It lists information like whether the query is a reversed one or not, and table_schema and query_schema versions.	2024-08-20 20:56:15 +02:00
Łukasz Paszkowski	727cbd8151	alternator::do_query: Use native reversed format When executing reversed queries, a native revered format shall be used. Therefore the table schema and the clustering key bounds are reversed before a partition slice and a read command are constructed. Similarly as for cql3::statements::select_statement.	2024-08-20 20:56:15 +02:00
Łukasz Paszkowski	3720e8aabe	alternator::do_query Rename schema with table_schema In order to increase readability, a schema variable is renamed to a table_schema to emphesize a table schema is passed to the function and used across it. Allows us to introduce a query_schema variable in the next patch.	2024-08-20 20:56:06 +02:00
Aleksandra Martyniuk	9d9414a75d	replica: add/remove table atomically Currently, database::tables_metadata::add_table needs to hold a write lock before adding a table. So, if we update other classes keeping track of tables before calling add_table, and the method yields, table's metadata will be inconsistent. Set all table-related info in tables_metadata::add_table_helper (called by add_table) so that the operation is atomic. Analogically for remove_table. Fixes: #19833. Closes scylladb/scylladb#20064	2024-08-20 20:53:32 +03:00
Kamil Braun	5c9efdff50	Merge 'raft: store_snapshot_descriptor to use actually preserved items number when truncating the local log table' from Sergey Zolotukhin io_fiber/store_snapshot_descriptor now gets the actual number of items preserved when the log is truncated, fixing extra entries remained after log snapshot creation. Also removes incorrect check for the number of truncated items in the raft_sys_table_storage::store_snapshot_descriptor. Minor change: Added error_injection test API for changing snapshot thresholds settings. Fixes scylladb/scylladb#16817 Fixes scylladb/scylladb#20080 Closes scylladb/scylladb#20095 * github.com:scylladb/scylladb: raft: Ensure const correctness in applier_fiber. raft: Invoke store_snapshot_descriptor with actually preserved items. raft: Use raft_server_set_snapshot_thresholds in tests. raft: Fix indentation in server.cc raft: Add a test to check log size after truncation. raft: Add raft_server_set_snapshot_thresholds injection. utils: Ensure const correctness of injection_handler::get().	2024-08-20 18:15:30 +02:00
Tomasz Grabiec	ff52527c54	Merge 'repair: do_rebuild_replace_with_repair: use source_dc only when safe' from Benny Halevy It is unsafe to restrict the sync nodes for repair to the source data center if it has too low replication factor in network_topology_replication_strategy, or if other nodes in that DC are ignored. Also, this change restricts the usage of source_dc to `network_topology` and `everywhere_topology` strategies, as with simple replication strategy there is no guarantee that there would be any more replicas in that data center. Fixes #16826 Reproducer submitted as https://github.com/scylladb/scylla-dtest/pull/3865 It fails without this fix and passes with it. * Requires backport to live versions. Issue hit in the filed with 2022.2.14 Closes scylladb/scylladb#16827 * github.com:scylladb/scylladb: repair: do_rebuild_replace_with_repair: use source_dc only when safe repair: replace_with_repair: pass the replace_node downstream repair: replace_with_repair: pass ignore_nodes as a set of host_id:s repair: replace_rebuild_with_repair: pass ks_erms from caller nodetool: rebuild: add force option Add and use utils::optional_param to pass source_dc	2024-08-20 16:13:23 +02:00
Sergey Zolotukhin	13b3d3a795	raft: Ensure const correctness in applier_fiber. Add 'const' to non mutable varibales in server_impl::applier_fiber() function.	2024-08-20 15:24:00 +02:00
Sergey Zolotukhin	c3e52ab942	raft: Invoke store_snapshot_descriptor with actually preserved items. - raft_sys_table_storage::store_snapshot_descriptor now receives a number of preserved items in the log, rather than _config.snapshot_trailing value; - Incorrect check for truncated number of items in store_snapshot_descriptor was removed. Fixes scylladb/scylladb#16817 Fixes scylladb/scylladb#20080	2024-08-20 15:22:49 +02:00
Sergey Zolotukhin	922e035629	raft: Use raft_server_set_snapshot_thresholds in tests. Replace raft_server_snapshot_reduce_threshold with raft_server_set_snapshot_thresholds in tests as raft_server_set_snapshot_thresholds fully covers the functionality of raft_server_snapshot_reduce_threshold.	2024-08-20 15:08:49 +02:00
Sergey Zolotukhin	00a1d3e305	raft: Fix indentation in server.cc	2024-08-20 15:08:45 +02:00
Sergey Zolotukhin	b6de8230a9	raft: Add a test to check log size after truncation. The test checks that snapshot_trailing_size parameter is taken into consideration when the log system table is truncated. Test for scylladb#16817	2024-08-20 14:15:50 +02:00
Sergey Zolotukhin	9dfa041fe1	raft: Add raft_server_set_snapshot_thresholds injection. Use error injection to allow overriding following snapshot threshold settings: - snapshot_threshold - snapshot_threshold_log_size - snapshot_trailing - snapshot_trailing_size	2024-08-20 14:15:50 +02:00
Sergey Zolotukhin	c5da0775f2	utils: Ensure const correctness of injection_handler::get(). Make utils::error_injection::injection_handler::get() method 'const' as it does not mutate object's state.	2024-08-20 14:15:50 +02:00
Botond Dénes	3ee0d7f2d1	Merge 'tools: Enhance scylla sstable shard-of to support tablets' from Kefu Chai before this change, `scylla sstable shard-of` didn't support tablets, because: - with tablets enabled, data distribution uses the scheduler - this replaces the previous method of mapping based on vnodes and shard numbers - as a result, we can no longer deduce sstable mapping from token ranges in this change, we: - read `system.tablets` table to retrieve tablet information - print the tablet's replica set (list of <host, shard> pairs) - this helps users determine where a given sstable is hosted This approach provides the closest equivalent functionality of `shard-of` in the tablet era. Fixes scylladb/scylladb#16488 --- no need to backport, it's an improvement, not a critical fix. Closes scylladb/scylladb#20002 * github.com:scylladb/scylladb: tools: enhance `scylla sstable shard-of` to support tablets replica/tablets: extract tablet_replica_set_from_cell() tools: extract get_table_directory() out tools: extract read_mutation out build: split the list of source file across multiple line tools/scylla-sstable: print warning when running shard-of with tablets	2024-08-20 13:51:12 +03:00
Avi Kivity	e2b179a3d0	Merge 'Coroutinize sstable_directory registry garbage collecting method' from Pavel Emelyanov null Closes scylladb/scylladb#20172 * github.com:scylladb/scylladb: sstable_directory: Coroutinize inner lambdas sstable_directory: Fix indentation after previous patch sstable_directory: Coroutinize outer cotinuation chain	2024-08-20 12:50:09 +03:00
David Garcia	fea707033f	docs: improve include flag directive The include flag directive now treats missing content as info logs instead of warnings. This prevents build failures when the enterprise-specific content isn't yet available. If the enterprise content is undefined, the directive automatically loads the open-source content. This ensures the end user has access to some content. address comments Closes scylladb/scylladb#19804	2024-08-20 12:21:39 +03:00
Kefu Chai	9a10c33734	build: cmake: do not build storage_proxy.o by default in `5ce07e5d84`, the target named "storage_proxy.o" was added for training the build of clang. but the rule for building this target has two flaws: * it was added a dependency of the "all" target, but we don't need to build `storage_proxy.cc` twice when building the tree in the regular build job. we only need to build it when creating the profile for training the build of clang. * it misses the include directory of abseil library. that's why we have following build failure when building the default target: ``` [2024-08-18T14:58:37.494Z] /usr/local/bin/clang++ -DDEBUG -DDEBUG_LSA_SANITIZER -DFMT_SHARED -DSANITIZE -DSCYLLA_BUILD_MODE=debug -DSCYLLA_ENABLE_ERROR_INJECTION -DSEASTAR_API_LEVEL=7 -DSEASTAR_DEBUG -DSEASTAR_DEBUG_PROMISE -DSEASTAR_DEBUG_SHARED_PTR -DSEASTAR_DEFAULT_ALLOCATOR -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SHUFFLE_TASK_QUEUE -DSEASTAR_SSTRING -DSEASTAR_TYPE_ERASE_MORE -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"Debug\" -I/jenkins/workspace/scylla-master/scylla-ci/scylla -I/jenkins/workspace/scylla-master/scylla-ci/scylla/seastar/include -I/jenkins/workspace/scylla-master/scylla-ci/scylla/build/seastar/gen/include -I/jenkins/workspace/scylla-master/scylla-ci/scylla/build/seastar/gen/src -I/jenkins/workspace/scylla-master/scylla-ci/scylla/build/gen -g -Og -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-enum-constexpr-conversion -Wno-unused-parameter -ffile-prefix-map=/jenkins/workspace/scylla-master/scylla-ci/scylla=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -U_FORTIFY_SOURCE -Werror=unused-result -fstack-clash-protection -fsanitize=address -fsanitize=undefined -fno-sanitize=vptr -MD -MT service/CMakeFiles/storage_proxy.o.dir/Debug/storage_proxy.cc.o -MF service/CMakeFiles/storage_proxy.o.dir/Debug/storage_proxy.cc.o.d -o service/CMakeFiles/storage_proxy.o.dir/Debug/storage_proxy.cc.o -c /jenkins/workspace/scylla-master/scylla-ci/scylla/service/storage_proxy.cc [2024-08-18T14:58:37.495Z] In file included from /jenkins/workspace/scylla-master/scylla-ci/scylla/service/storage_proxy.cc:17: [2024-08-18T14:58:37.495Z] In file included from /jenkins/workspace/scylla-master/scylla-ci/scylla/db/commitlog/commitlog.hh:19: [2024-08-18T14:58:37.495Z] In file included from /jenkins/workspace/scylla-master/scylla-ci/scylla/db/commitlog/commitlog_entry.hh:15: [2024-08-18T14:58:37.495Z] In file included from /jenkins/workspace/scylla-master/scylla-ci/scylla/mutation/frozen_mutation.hh:15: [2024-08-18T14:58:37.495Z] In file included from /jenkins/workspace/scylla-master/scylla-ci/scylla/mutation/mutation_partition_view.hh:16: [2024-08-18T14:58:37.495Z] In file included from /jenkins/workspace/scylla-master/scylla-ci/scylla/build/gen/idl/mutation.dist.impl.hh:14: [2024-08-18T14:58:37.495Z] /jenkins/workspace/scylla-master/scylla-ci/scylla/serializer_impl.hh:20:10: fatal error: 'absl/container/btree_set.h' file not found [2024-08-18T14:58:37.495Z] 20 \| #include <absl/container/btree_set.h> [2024-08-18T14:58:37.495Z] \| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~ [2024-08-18T14:58:37.495Z] 1 error generated. ``` * if user only enables "dev" mode, we'd have: ``` CMake Error at service/CMakeLists.txt:54 (add_library): No SOURCES given to target: storage_proxy.o ``` so, in this change, we * exclude this target from "all" * link this target against abseil header library, so it has access to the abseil library. please note, we don't need to build an executable in this case, so the header would suffice. * add a proxy target to conditionally enable/disable this target. as CMake does not support generator expression in `add_dependencies()` yet at the time of writing. see https://gitlab.kitware.com/cmake/cmake/-/issues/19467 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20195	2024-08-19 21:30:34 +03:00
Avi Kivity	7eb3b15fff	Merge 'utils/tagged_integer: remove conversion to underlying integer' from Laszlo Ersek ~~~ utils/tagged_integer: remove conversion to underlying integer Silently converting a tagged (i.e., "dimension-ful") integer to a naked ("dimensionless") integer defeats the purpose of having tagged integers, and is a source of practical bugs, such as <https://github.com/scylladb/scylladb/issues/20080>. We could make the conversion operator explicit, for enforcing static_cast<TAGGED_INTEGER_TYPE::value_type>(TAGGED_INTEGER_VALUE) in every conversion location -- but that's a mouthful to write. Instead, remove the conversion operator, and let clients call the (identically behaving) value() member function. ~~~ No backport needed (refactoring). The series is supposed to solve #20081. Two patches in the series touch up code that is known to be (orthogonally) buggy; see - `service/raft_sys_table_storage: tweak dead code` (#20080) - `test/raft/replication: untag index_t in test_case::get_first_val()` (#20151) Fixes for those (independent) issues will have to be rebased on this series, or this series will have to be rebased on those (due to context conflicts). The series builds at every stage. The debug and release unit test suites pass at the end. Closes scylladb/scylladb#20159 * github.com:scylladb/scylladb: utils/tagged_integer: remove conversion to underlying integer test/raft/randomized_nemesis_test: clean up remaining index_t usage test/raft/randomized_nemesis_test: clean up index_t usage in store_snapshot() test/raft/replication: clean up remaining index_t usage test/raft/replication: take an "index_t start_idx" in create_log() test/raft/replication: untag index_t in test_case::get_first_val() test/raft/etcd_test: tag index_t and term_t for comparisons and subtractions test/raft/fsm_test: tag index_t and term_t for comparisons and subtractions test/raft/helpers: tighten compare_log_entries() param types service/raft_sys_table_storage: tweak dead code service/raft_sys_table_storage: simplify (snap.idx - preserve_log_entries) service/raft_sys_table_storage: untag index_t and term_t for queries raft/server: clean up index_t usage raft/tracker: don't drop out of index_t space for subtraction raft/fsm: clean up index_t and term_t usage raft/log: clean up index_t usage db/system_keyspace: promise a tagged integer from increment_and_get_generation() gms/gossiper: return "strong_ordering" from compare_endpoint_startup() gms/gossiper: get "int32_t" value of "gms::version_type" explicitly	2024-08-19 19:52:54 +03:00
Benny Halevy	5f655e41e3	repair: do_rebuild_replace_with_repair: use source_dc only when safe It is unsafe to restrict the sync nodes for repair to the source data center if we cannot guarantee a quorum in the data center with network-topology replication strategy. This change restricts the usage of source_dc in the following cases: 1. For SimpleStrategy - source_dc is ignored since there is no guarantee that it contains remaining replicas for all tokens. 2. For EverywhereStrategy - use source_dc if there are remaining live nodes in the datacenter. 3. For NetworkTopologyStrategy: a. It is considered unsafe to use source_dc if number of nodes lost in that DC (replaced/rebuilt node + additional ignored nodes) is greater than 1, or it has 1 lost node and rf <= 1 in the DC. b. If the source_dc arg is forced, as with the new `nodetool rebuild --force <source_dc>` option, we use it anyway, even if it's considered to be unsafe. A warning is printed in this case. c. If the source_dc arg is user-provided, (using nodetool rebuild), an error exception is thrown, advising to use an alternative dc, if available, omit source_dc to sync with all nodes, or use the --force option to use the given source_dc anyhow. d. Otherwise, we look for an alternative source datacenter, that has not lost any node. If such datacenter is found we use it as source_dc for the keyspace, and log a warning. e. If no alternative dc is found (and source_dc is implicit), then: log a warning and fall back to using replicas from all nodes in the cluster. Fixes #16826 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-19 17:23:51 +03:00
Benny Halevy	8665eef98c	repair: replace_with_repair: pass the replace_node downstream To be used by the next path to count how many nodes are lost in each datacenter. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-19 17:23:33 +03:00
Benny Halevy	9729dd21c3	repair: replace_with_repair: pass ignore_nodes as a set of host_id:s The callers already pass ignore_nodes as host_id:s and we translate them into inet_address only for repair so delay the translation as much as posible, Refs scylladb/scylladb#6403 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-19 17:22:01 +03:00
Benny Halevy	b5d0ab092c	repair: replace_rebuild_with_repair: pass ks_erms from caller The keyspaces replication maps must be in sync with the token_metadata_ptr passed already to the functions, so instead of getting it in the callee, let the caller get the ks_erms along with retrieving the tmptr. Note that it's already done on the rebuild path for streaming based rebuild. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-19 17:20:27 +03:00
Benny Halevy	0419b1d522	nodetool: rebuild: add force option To be used to force usage of source_dc, even when it is unsafe for rebuild. Update docs and add test/nodetool/test_rebuild.py Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-19 17:20:12 +03:00
Benny Halevy	8b1877f3ca	Add and use utils::optional_param to pass source_dc Clearly indicate if a source_dc is provided, and if so, was it explicitly given by the user, or was implicitly selected by scylla. This will become useful in the next patches that will use that to either reject the operation if it's unsafe to use the source_dc and the dc was explicitly given by the user, or whether to fallback to using all nodes otherwise. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-19 17:13:54 +03:00
Anna Stuchlik	83d5cb04c2	doc: extract the info about tablets defaut to a separate file This commit extracts the information about the default for tables in keyspace creation to a separate file in the _common folder. The file is then included using the scylladb_include_flag directive. The purpose of this commit is to make it possible to include a different file in the scylla-enterprise repo - with a different default. Refs https://github.com/scylladb/scylla-enterprise/issues/4585 Closes scylladb/scylladb#20181	2024-08-19 16:16:18 +03:00
Kefu Chai	25b3c50f71	test/nodetool: print default value of options in help message would be more helpful, if the output of "--help" command line can include the default value of options. so, in this change, we include the default values in it. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20170	2024-08-19 16:15:24 +03:00
Botond Dénes	40d2a6f0b2	Merge 'test.py: use XPath for iterating in "TestSuite/TestSuite"' from Kefu Chai before this change, we check for the existence of "TestSuite" node under the root of XML tree, and then enumerating all "TestSuite" nodes under this "TestSuite", this approach works. but it * introduces unnecessary indent * is not very readable in this change, we just use "./TestSuite/TestSuite" for enumerating all "TestSuite" nodes under "TestSuite". simpler this way. --- it's a cleanup in the test driver script, hence no need to backport. Closes scylladb/scylladb#20169 * github.com:scylladb/scylladb: test.py: fix the indent test.py: use XPath for iterating in "TestSuite/TestSuite"	2024-08-19 16:13:42 +03:00
Botond Dénes	6835f7e993	Merge 'Add CQL-based RBAC support to Alternator' from Piotr Smaron Alternator already supports authentication - the ability to to sign each request as a particular user. The users that can be used are the different "roles" that are created by CQL "CREATE ROLE" commands. This series adds support for authorization, i.e., the ability to determine that only some of these roles are allowed to read or write particular tables, to create new tables, and so on. The way we chose to do this in this series is to support CQL's existing role-based access control (RBAC) commands - GRANT and REVOKE - on Alternator tables. For example, an Alternator table "xyz" is visible to CQL as "alternator_xyz.xyz", so a `GRANT SELECT ON alternator_xyz.xyz TO myrole` will allow read commands (e.g., GetItem) on that table, and without this GRANT, a GetItem will fail with `AccessDeniedException`. This series adds the necessary checks to all relevant Alternator operations, and also adds extensive functional testing for this feature - i.e., that certain DynamoDB API operations are not allowed without the appropriate GRANTs. The following permissions are needed for the following Alternator API operations: * SELECT: `GetItem`, `Query`, `Scan`, `BatchGetItem`, `GetRecords` * MODIFY: `PutItem`, `DeleteItem`, `UpdateItem`, `BatchWriteItem` * CREATE: `CreateTable` * DROP: `DeleteTable` * ALTER: `UpdateTable`, `TagResource`, `UntagResource`, `UpdateTimeToLive` * _none needed_: `ListTables`, `DescribeTable`, `DescribeEndpoints`, `ListTagsOfResource`, `DescribeTimeToLive`, `DescribeContinuousBackups`, `ListStreams`, `DescribeStream`, `GetShardIterator` Currently, I decided that for consistency each operation requires one permission only. For example, PutItem only requires MODIFY permission. This is despite the fact that in some cases (namely, `ReturnValues=ALL_OLD`) it can also _read_ the item. We should perhaps discuss this decision - and compare how it was done in CQL - e.g., what happens in LWT writes that may return old values? Different permissions can be granted for a base table, each of its views, and the CDC table (Alternator streams). This adds power - e.g., we can allow a role to read only a view but not the base table, or read the table but not its history. GRANTing permissions on views or CDC logs require knowing their names, which are somewhat ugly (e.g., the name of GSI "abc" in table "xyz" is `alternator_xyz.xyz:abc`). But usefully, the error message when permissions are denied contains the full name of the table that was lacking permissions and which permissions were lacking, so users can easily add them. In addition to permissions checking, this series also correctly supports _auto-grant_ (except #19798): When a role has permissions to `CreateTable`, any table it creates will automatically be granted all permissions for this role, so this role will be able to use the new table and eventually delete it. `DeleteTable` does the opposite - it removes permissions from tables being deleted, so that if later a second user re-creates a table with the same name, the first user will not have permissions over the new table. The already-existing configuration parameter `alternator_enforce_authorization` (off by default), which previously only enabled authentication, now also enables authorization. Users that upgrade to the new version and already had `alternator_enforce_authorization=true` should verify that the users they use to authenticate either have the appropriate permissions or the "superuser" flag. Roles used to authenticate must also have the "login" flag. Please note that although the new RBAC support implements the access control feature we asked for in #5047, this implementation is _not compatible_ with DynamoDB. In DynamoDB, the access control is configured through IAM operations or through the new `PutResourcePolicy` - operation, not through CQL (obviously!). DynamoDB also offers finer access-control granularity than we support (Scylla's RBAC works on entire tables, DynamoDB allows setting permissions on key prefixes, on individual attributes, and more). Despite this non-compatibility, I believe this feature, as is, will already be useful to Alternator users. Fixes #5047 (after closing that issue, a new clean issue should be opened about the DynamoDB-compatible APIs that we didn't do - just so we remember this wasn't done yet). New feature, should not be backported. Closes scylladb/scylladb#20135 * github.com:scylladb/scylladb: tests: disable test_alternator_enforce_authorization_true test, alternator: test for alternator_enforce_authorization config test/pylib: allow setting driver_connect() options in servers_add() test: fix test_localnodes_joining_nodes alternator, RBAC: reproducer for missing CDC auto-grant alternator: document the new RBAC support alternator: add RBAC enforcement to GetRecords test/alternator: additional tests for RBAC test/alternator: reduce permissions-validity-in-ms test/alternator: add test for BatchGetItem from multiple tables alternator: test for operations that do not need any permissions alternator: add RBAC enforcement to UpdateTimeToLive alternator: add RBAC enforcement to TagResource and UntagResource alternator: add RBAC enforcement to BatchGetItem alternator: add RBAC enforcement to BatchWriteItem alternator: add RBAC enforcement to UpdateTable alternator: add RBAC enforcement to Query and Scan alternator: add RBAC enforcement to CreateTable alternator: add RBAC enforcement to DeleteTable alternator: add RBAC enforcement to UpdateItem alternator: add RBAC enforcement to DeleteItem alternator: add RBAC enforcement to PutItem alternator: add RBAC enforcement to GetItem alternator: stop using an "internal" client_state	2024-08-19 16:09:53 +03:00
Tomasz Grabiec	c1de4859d8	Merge 'tablets: Fix race between repair and split' from Raphael "Raph" Carvalho Consider the following: ``` T 0 split prepare starts 1 repair starts 2 split prepare finishes 3 repair adds unsplit sstables 4 repair ends 5 split executes ``` If repair produces sstable after split prepare phase, the replica will not split that sstable later, as prepare phase is considered completed already. That causes split execution to fail as replicas weren't really prepared. This also can be triggered with load-and-stream which shares the same write (consumer) path. The approach to fix this is the same employed to prevent a race between split and migration. If migration happens during prepare phase, it can happen source misses the split request, but the tablet will still be split on the destination (if needed). Similarly, the repair writer becomes responsible for splitting the data if underlying table is in split mode. That's implemented in replica::table for correctness, so if node crashes, the new sstable missing split is still split before added to the set. Fixes #19378. Fixes #19416. *Please replace this line with justification for the backport/\ labels added to this PR** Closes scylladb/scylladb#19427 * github.com:scylladb/scylladb: tablets: Fix race between repair and split compaction: Allow "offline" sstable to be split	2024-08-19 14:44:28 +02:00
Kefu Chai	151074240c	utils: cached_file: use structured binding when appropriate for better readability. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20184	2024-08-19 14:01:42 +03:00
Piotr Smaron	f773c76bfb	codeowners: add appropriate reviewers to the cluster components	2024-08-19 12:39:47 +02:00
Anna Stuchlik	8fb746a5d2	doc: fix a link on the RBAC page This commit fixes an external link on the Role Based Access Control page. Fixes https://github.com/scylladb/scylladb/issues/20166 Closes scylladb/scylladb#20171	2024-08-19 12:56:38 +03:00
Piotr Smaron	cdc88cd06c	tests: disable test_alternator_enforce_authorization_true The test is flaky and needs to be fixed in order to not randomly break our CI, OTOH can be commented out for the time being, so that we can marge the feature.	2024-08-19 09:57:53 +02:00
Nadav Har'El	989dbef315	test, alternator: test for alternator_enforce_authorization config This patch adds tests that demonstrates the current way that Alternator's authentication and authorization are both enabled or disabled by the option "alternator_enforce_authorization". If in the future we decide to change this option or eliminate it (e.g., remain just with the "authenticator" and "authorizer" options), we can easily update these tests to fit the new configuration parameters and check they work as expected. Because the new tests want to start Scylla instances with different configuration parameters, they are written in the the "topology" framework and not in the test/alternator framework. The test/alternator framework still contains (test/alternator/test_cql_rbac.py) the vast majority of the functional testing of the RBAC feature where all those tests just assume that RBAC is enabled and needs to be tested. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	41418603e1	test/pylib: allow setting driver_connect() options in servers_add() The manager.driver_connect() functions allows to pass parameters when creating the connection (e.g., a special auth_provider), but unfortunately right now the servers_add() function always calls driver_connect() without parameters. So in this patch we just add a new optional parameter to servers_add(), driver_connect_opts, that will be passed to driver_connect(). In theory instead of the new option to driver_connect() a caller can pass start=False to servers_add() and later call driver_connect() manually with the right arguments. The problem is that start=False avoids more than just calling driver_connect(), so it doesn't solve the problem. An example of using the new option is to run Scylla with authentication enabled, and then connect to it using the correct default account ("cassandra"/"cassandra"): config = { 'authenticator': 'PasswordAuthenticator', 'authorizer': 'CassandraAuthorizer' } servers = await manager.servers_add(1, config=config, driver_connect_opts={'auth_provider': PlainTextAuthProvider(username='cassandra', password='cassandra')}) Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	de20ac1a6d	test: fix test_localnodes_joining_nodes The existing test topology_experimental_raft/test_alternator::test_localnodes_joining_nodes Tried to create a second server but not wait for it to complete, but the trick it used (cancelling the task) doesn't work since commit `2ee063c` makes a list of unwaited tasks and waits for them anyway. The test appears to work because it is the last test in the file, but if we ever add another test in the same file (like I plan to do in the next patch), that other test will find a "BROKEN" ScyllaClusterManager and report that it failed :-( Other tricks I tried to use (like killing the servers) also didn't work because of various limitations and complications of the test framework and all its layers. So not wanting to fight the fragile testing framework any more at this point, I just gave up and the test will wait for the second server to come up. This adds 120 seconds (!) to the test, but since this whole test file already takes more than 500 seconds to complete, let's bite this bullet. Maybe in the future when the test framework improves, we can avoid this 120 second wait. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	79f9b3007e	alternator, RBAC: reproducer for missing CDC auto-grant This patch adds a reproducing (xfailing) test for issue #19798, which shows that if a role is able to create an Alternator table, the role is able to read the new table (this is known as "auto-grant"), but is NOT able to read the CDC log (i.e., use Alternator Streams' "GetRecords"). Once we do fix this auto-grant bug, it's also important to also implement auto-revoke - the permissions on a deleted table must be deleted as well (otherwise the old owner of a deleted table will be able to read a new table with the same name). This patch also adds a test verifying that auto-revoke works. This test currently passes (because there is no auto- grant, so nothing needs to be revoked...) but if we'll implement auto-grant and forget auto-revoke, the second test will start to fail - so I added this test as a precaution against a bad fix. Refs #19798 Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	7de6aedd47	alternator: document the new RBAC support In docs/alternator/compatibility.md we said that although Alternator supports authentication, it doesn't support authorization (access control). Now it does, so the relevant text needs to be corrected to fit what we have today. It's still in the compatibility.md document because it's not the same API as DynamoDB's, so users with existing applications may need to be aware of this difference. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	f9ff475dfb	alternator: add RBAC enforcement to GetRecords This patch adds a requirement for the "SELECT" permission on a table to run a GetRecords on it (the DynamoDB Streams API, i.e., CDC). The grant is checked on the CDC log table - not on the base table, which allows giving a role the ability to read the base but not is change stream, or vice versa. The operations ListStreams, DescribeStreams, GetShardIterators do not require any permissions to run - they do not read any data, and are (in my opinion) similar in spirit to DescribeTable, so I think it's fine not to require any permissions for them. A test is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	0789841cf8	test/alternator: additional tests for RBAC Additional tests for support for CQL Role-Based Access Control (RBAC) in Alternator: 1. Check that even in an Alternator table whose name isn't valid as CQL table names (e.g., uses the dot character) the GRANT/REVOKE commands work as expected. 2. Check that superuser roles have full permissions, as expected. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	409fea5541	test/alternator: reduce permissions-validity-in-ms We set in test/cql-pytest/run.py, affecting test/alternator/run, the configuration permissions_validity_in_ms by default to 100ms. This means that tests that need to check how GRANT or REVOKE work always need to sleep for more than 100ms, which can make a test with a lot of these operations very slow. So let's just set this configuration value to 5ms. I checked that it doesn't adversely affect the total running speed of test/alternator/run. This change only affects running tests through test/alternator/run, which is expected to be fast. I left the default for test.py as it was, 100ms, the latency of individual tests is less important there. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	1b20a11dec	test/alternator: add test for BatchGetItem from multiple tables While working on the RBAC on BatchGetItem, I noticed that although BatchGetItem may ask to read items from several tables, we don't have a test covering this case! This patch fixes that testing oversight. Note that for the write-side version of this operation, BatchWriteItem, we do have tests that write to several tables in the same batch. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	f827bd51d2	alternator: test for operations that do not need any permissions Some operations, namely ListTables, DescribeTable, DescribeEndpoints, ListTagsOfResource, DescribeTimeToLive and DescribeContinuousBackups do not need any permissions to be GRANTed to a role. Our rationale for this decision is that in CQL, "describe table" and friends also do not require any permissions. This patch includes a test that verifies that they really don't need permissions. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	9417cf8bcf	alternator: add RBAC enforcement to UpdateTimeToLive This patch adds a requirement for the "ALTER" permission on a table to run a UpdateTimeToLive on it. UpdateTimeToLive is similar in purpose to UpdateTable, so it makes sense to use the same permission "ALTER" as we do for UpdateTable. A tests is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	e76316495c	alternator: add RBAC enforcement to TagResource and UntagResource This patch adds a requirement for the "ALTER" permission on a table to run the TagResource or UntagResource operations on it. CQL does not have an exact parallel of DynamoDB's tagging feature, but our usual use of tags as an extension of UpdateTable to change non-standard options (e.g., write isolation policy or tablets setup), so it makes sense to require the same permissions we require for UpdateTable - namely "ALTER". A test for both operations is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:53 +02:00
Nadav Har'El	fda4a9fad8	alternator: add RBAC enforcement to BatchGetItem This patch adds a requirement for the "SELECT" permission on a table to run a BatchGetItem on it. A single batch may ask to write to several different tables, so we fail the entire batch with AccessDeniedException if any of the tables mentioned in the batch do not have SELECT permissions for this role. A tests is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:57:51 +02:00
Nadav Har'El	b02288785f	alternator: add RBAC enforcement to BatchWriteItem This patch adds a requirement for the "MODIFY" permission on a table to run a BatchWriteItem on it. A single batch may ask to write to several different tables, so we fail the entire batch with AccessDeniedException if any of the tables mentioned in the batch do not have MODIFY permissions for this role. A tests is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:56:28 +02:00
Nadav Har'El	445a5d57cd	alternator: add RBAC enforcement to UpdateTable This patch adds a requirement for the "ALTER" permission on a table to run a UpdateTable on it. A tests is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	b4484158e7	alternator: add RBAC enforcement to Query and Scan This patch adds a requirement for the "SELECT" permission on a table to run a Query or Scan on it. Both Query and Scan operations call the same do_query() function, so the permission checks are put there. Note that Query can read from either the base table or one of its views, and the permissions on the base and each of the views can be separate (so we can allow a role to only read one view, for example). Tests for all of the above are also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	82f7e55943	alternator: add RBAC enforcement to CreateTable This patch adds a requirement for the "CREATE" permission on ALL KEYSPACES to run a CreateTable operation. The CreateTable operation also performs so-called "auto-grant": When a role creates a table, it is automatically granted full permissions to read, write, change or delete that new table. A test for all these things is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	79dfb7b7d5	alternator: add RBAC enforcement to DeleteTable This patch adds a requirement for the "DROP" permission on a table to run a DeleteTable on it. Moreover, when a table and its views are deleted, any special permissions previously GRANTed on this table are removed. This is necessary because if a role creates a table it is automatically granted permissions on this table (this is known as "auto-grant" - see the CreateTable patch for details). If this role deletes this table and later a second role creates a table with the same name, we don't want the first role to have permissions on this new table. Tests for permission enforcements and revocation on delete are also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	2ebc0501b8	alternator: add RBAC enforcement to UpdateItem This patch adds a requirement for the "MODIFY" permission on a table to run a UpdateItem on it. Only the MODIFY permission is required, even if the operation may also read the old value of the item, such as a read-modify-write operation or even using ReturnValues='ALL_OLD'. A test is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	36d8aea654	alternator: add RBAC enforcement to DeleteItem This patch adds a requirement for the "MODIFY" permission on a table to run a DeleteItem on it. Only the MODIFY permission is required, even if the operation may also read the old value of the item (using ReturnValues='ALL_OLD'). A test is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	34c975854a	alternator: add RBAC enforcement to PutItem This patch adds a requirement for the "MODIFY" permission on a table to run a PutItem on it. Only the MODIFY permission is required, even if the operation may also read the old value of the item (using ReturnValues='ALL_OLD'). A test is also added. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	3008b8416c	alternator: add RBAC enforcement to GetItem In this patch, we begin to add role-based access control (RBAC) enforement to Alternator - in this patch only to GetItem. After the preparation of client_state correctly in the previous patch, the permission check itself in the get_item() function is very simple. The bigger part of this patch is a full functional test in test/alternator/test_cql_rbac.py. The test is quite self-explanatory and heavily commented. Basically we check that a new role cannot read with GetItem a pre-existing table, and we can add that ability by GRANTing (in CQL) the new role the ability to SELECT the table, the keyspace, all keyspaces, or add that ability to some other role that this role inherits. In the following patches, we will add role-based access control to the Alternator operations, but the functional tests will be shorter - we don't need to check the role inheritence, "all keyspaces" feature, and so on, for every operation separately since they all use the same underlying checking functions which handles these role inheritence issues in exactly the same way. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Nadav Har'El	583f060bd8	alternator: stop using an "internal" client_state Scylla uses a "client_state" object to encapsulate the information of who the client is - its IP address, which user was authenticated, and so on. For an unknown reason, Alternator created for each request an "internal" client_state, meaning that supposedly the client for each request was some sort of internal process (e.g., repair) rather than a real client. This was wrong, and we even had a FIXME about not putting the client's IP address in client_state. So in this patch, we start using a normal "external" client_state instead of an "internal" one. The client_state constructors are very different in the two cases, so a few lines of code had to change. I hope that this change will cause no functional changes. For example, Alternator was already setting its own timeouts explicitly and not relying on the default ones for external clients. However, we need to fix this for the following patches which introduce permissions checks (Role-Based Access Control - RBAC) - the client_state methods for checking permissions become no-ops for internal clients (even if the client_state contains an authenticated users). We need these functions to do their job - so we need an external variant of client_state. Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2024-08-19 09:45:22 +02:00
Tomasz Grabiec	ab7656a7be	Merge 'replica: fix copy constructor of tablet_sstable_set' from Lakshmi Narayanan Sreethar Commit `9f93dd9fa3` changed `tablet_sstable_set::_sstable_sets` to be a `absl::flat_hash_map` and in addition, `std::set<size_t> _sstable_set_ids` was added. `_sstable_set_ids` is set up in the `tablet_sstable_set(schema_ptr s, const storage_group_manager& sgm, const locator::tablet_map& tmap)` constructor, but it is not copied in `tablet_sstable_set(const tablet_sstable_set& o)`. This affects the `tablet_sstable_set::tablet_sstable_set` method as it depends on the copy constructor. Since sstable set can be cloned when a new sstable set is added, the issue will cause ids not being copied into the new sstable set. It's healed only after compaction, since the sstable set is rebuilt from scratch there. This PR fixes this issue by removing the existing copy constructor of `tablet_sstable_set` to enable the implicit default copy constructor. Fixes #19519 Closes scylladb/scylladb#20115 * github.com:scylladb/scylladb: boost/sstable_set_test: add testcase to test tablet_sstable_set copy constructor replica: fix copy constructor of tablet_sstable_set	2024-08-19 00:53:29 +02:00
Avi Kivity	390e01673b	Merge 'Adding batch latency and batch size metrics to Alternator' from Amnon Heiman This patch adds metrics for batch get_item and batch write_item. The new metrics record summary and histogram for latencies and batch size. Batch sizes are implemented as ever-growing counters. To get the average batch size divide the rate of the batch size counter by the rate of the number of batch counter: ```rate(batch_get_item_batch_size)/rate(batch_get_item)``` Relates to #17615 New code, No need to backport Closes scylladb/scylladb#20190 * github.com:scylladb/scylladb: Add tests for Alternator batch operation metrics alternator/executor: support batch latency and size metrics Add metrics for Alternator get and write batch operations	2024-08-18 21:22:39 +03:00
Amnon Heiman	63fdfb89cd	Add tests for Alternator batch operation metrics This patch adds unit tests to verify the correctness of the newly introduced histogram metrics for get and write batch operation latencies. The test uses the existing latency test with the added metrics. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-08-18 12:19:43 +03:00
Amnon Heiman	d20a333f51	alternator/executor: support batch latency and size metrics This patch Updated the get and write batch operations in Alternator to record latency using the newly added histogram metrics. It adds logic to increment the counters with the number of items processed in each batch. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-08-18 12:14:23 +03:00
Amnon Heiman	8bad4b44f8	Add metrics for Alternator get and write batch operations Introduced histogram metrics to track latency for Alternator's get and write batch operations. Added counters to record the number of items processed in each batch operation. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2024-08-18 12:09:46 +03:00
Lakshmi Narayanan Sreethar	ec47b50859	boost/sstable_set_test: add testcase to test tablet_sstable_set copy constructor Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-08-17 23:38:05 +05:30
Lakshmi Narayanan Sreethar	44583eed9e	replica: fix copy constructor of tablet_sstable_set Remove the existing copy constructor to enable the use of the implicit copy constructor. This fixes the issue of `_sstable_set_ids` not being copied in the current copy constructor. Fixes #19519 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-08-17 23:37:58 +05:30
Kefu Chai	3d593ceeb1	perf/perf_sstable: add {crawling,partitioned}_streaming modes for testing the load performance of load_and_stream operation. Refs scylladb/scylladb#19989 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-17 14:43:54 +08:00
Kefu Chai	7806c72e49	test/perf/perf_sstable: use switch-case when appropriate this change is a follow up of `06c60f6ab`, which updated the 2nd step of the test to use switch-case, but missed the 1st step. so this change updates the first step of the test as well. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-17 14:38:37 +08:00
Pavel Emelyanov	6a9b8ea135	sstable_directory: Coroutinize inner lambdas Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-16 10:45:27 +03:00
Pavel Emelyanov	7401c0ace2	sstable_directory: Fix indentation after previous patch Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-16 10:45:27 +03:00
Pavel Emelyanov	7422504d35	sstable_directory: Coroutinize outer cotinuation chain Indentation is deliberately left broken Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-16 10:45:27 +03:00
Kefu Chai	e8f9f71ef3	test.py: fix the indent and take this opportunity to fix a typo in comment. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-16 13:32:57 +08:00
Kefu Chai	e88166f7a4	test.py: use XPath for iterating in "TestSuite/TestSuite" before this change, we check for the existence of "TestSuite" node under the root of XML tree, and then enumerating all "TestSuite" nodes under this "TestSuite", this approach works. but it * introduces unnecessary indent * is not very readable in this change, we just use "./TestSuite/TestSuite" for enumerating all "TestSuite" nodes under "TestSuite". simpler this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-16 13:32:33 +08:00
Kefu Chai	afee3924b3	s3/client: check for "Key" and "Value" tag in "Tag" XML tag despite that the API document at https://docs.aws.amazon.com/AmazonS3/latest/API/API_Tag.htm claims that both these tags are "Required" in the "Tag" object returned by S3 APIs, we still have to check them before dereferencing the pointer of the child node, as we should not trust the output of an external API. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20160	2024-08-15 20:16:35 +03:00
Andrei Chekun	f24f5b7db2	test.py: Fix boost XML conversion to allure when XML file is empty The method cannot find the TestSuite in the XML file and fails the whole job, however tests are passed. The issue was in incorrect understanding of boost summarization method. It creates one file for all modes, so there is no need to go through all modes to convert the XML file for allure. Closes: https://github.com/scylladb/scylladb/issues/20161 Closes scylladb/scylladb#20165	2024-08-15 20:15:31 +03:00
Benny Halevy	52234214e5	schema_tables: calculate_schema_digest: filter the key earlier Currently, each frozen mutation we get from system_keyspace::query_mutations is unfrozen in whole to a mutation and only then we check its key with the provided `accept_keyspace` function. This is wasteful, since they key can be processed directly form the frozen mutation, before taking the toll of unfreezing it. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-15 12:33:34 +03:00
Benny Halevy	95a5fba0ea	schema_tables: calculate_schema_digest: prevent stalls due to large mutations vector With a large number of table the schema mutations vector might get big enoug to cause reactor stalls when freed. For example, the following stall was hit on 2023.1.0~rc1-20230208.fe3cc281ec73 with 5000 tables: ``` (inlined by) ~vector at /usr/bin/../lib/gcc/x86_64-redhat-linux/12/../../../../include/c++/12/bits/stl_vector.h:730 (inlined by) db::schema_tables::calculate_schema_digest(seastar::sharded<service::storage_proxy>&, enum_set<super_enum<db::schema_feature, (db::schema_feature)0, (db::schema_feature)1, (db::schema_feature)2, (db::schema_feature)3, (db::schema_feature)4, (db::schema_feature)5, (db::schema_feature)6, (db::schema_feature)7> >, seastar::noncopyable_function<bool (std::basic_string_view<char, std::char_traits<char> >)>) at ./db/schema_tables.cc:799 ``` This change returns a mutations generator from the `map` lambda coroutine so we can process them one at a time, destroy the mutations one at a time, and by that, reducing memory footprint and preventing reactor stalls. Fixes #18173 Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-08-15 12:33:34 +03:00
Kefu Chai	c628fa4e9e	tools: enhance `scylla sstable shard-of` to support tablets before this change, `scylla sstable shard-of` didn't support tablets, because: - with tablets enabled, data distribution uses the scheduler - this replaces the previous method of mapping based on vnodes and shard numbers - as a result, we can no longer deduce sstable mapping from token ranges in this change, we: - read `system.tablets` table to retrieve tablet information - print the tablet's replica set (list of <host, shard> pairs) - this helps users determine where a given sstable is hosted This approach provides the closest equivalent functionality of `shard-of` in the tablet era. Fixes scylladb/scylladb#16488 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-15 15:49:55 +08:00
Kefu Chai	4291033b14	replica/tablets: extract tablet_replica_set_from_cell() so it can be reused to implement a low-level tool which reads tablets data from sstables Refs scylladb/scylladb#16488 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-15 15:49:55 +08:00
Kefu Chai	e1162e0dae	tools: extract get_table_directory() out the `get_table_directory()` function will have applications beyond its current use in `schema_loader.cc`. its ability to locate the directory storing the sstables of given table could be valuable in other subcommand(s) implementation. so, in this change we extract it out into a dedicated source file, so that it accept the primary_key and an optional clustering_key. Refs scylladb/scylladb#16488 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-15 15:49:55 +08:00
Kefu Chai	a04e0b6c7d	tools: extract read_mutation out the `read_mutation_from_table_offline()` function will have applications beyond its current use in `schema_loader.cc`. its ability to parser mutation data from sstables could be valuable in other subcommand(s) implementation. so, in this change we extract it out into a dedicated source file, so that it accept the primary_key and an optional clustering_key. Refs scylladb/scylladb#16488 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-15 15:49:55 +08:00
Kefu Chai	74a670dd19	build: split the list of source file across multiple line Split the extended list of source files across multiple lines. This improves readability and makes future additions easier to review in diffs. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-15 15:49:55 +08:00
Kefu Chai	3f8f1d7274	tools/scylla-sstable: print warning when running shard-of with tablets the subcommand of "shard-of" does not support tablets yet. so let's print out an error message, instead of printing the mapping assuming that the sstables are distributed based on token only. this commit also adds two more command line options to this subcommand, so that user is required to specify either "--vnodes" or "--tablets" to instruct the tool how the cluster distributes the tokens across nodes and their shards. this helps to minimize the suprise of user. this change prepares for the succeeding changes to implement the tablets support. the corresponding test is updated accordingly so that it only exercises the "shard-of" subcommand without tablets. we will test it with tablets enabled in a succeeding change. Refs scylladb/scylladb#16488 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-15 15:49:55 +08:00
Laszlo Ersek	baf6ec49ff	utils/tagged_integer: remove conversion to underlying integer Silently converting a tagged (i.e., "dimension-ful") integer to a naked ("dimensionless") integer defeats the purpose of having tagged integers, and is a source of practical bugs, such as <https://github.com/scylladb/scylladb/issues/20080>. We could make the conversion operator explicit, for enforcing static_cast<TAGGED_INTEGER_TYPE::value_type>(TAGGED_INTEGER_VALUE) in every conversion location -- but that's a mouthful to write. Instead, remove the conversion operator, and let clients call the (identically behaving) value() member function. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-15 02:12:58 +02:00
Laszlo Ersek	9aa7d232d6	test/raft/randomized_nemesis_test: clean up remaining index_t usage With implicit conversion of tagged integers to untagged ones going away, explicitly tag (or untag, as necessary) the operands of the following operations, in "test/raft/randomized_nemesis_test.cc": - addition of tagged and untagged (both should be tagged) - taking the minimum of an index difference and a container size (both should be untagged) Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Laszlo Ersek	1af3460a81	test/raft/randomized_nemesis_test: clean up index_t usage in store_snapshot() With implicit conversion of tagged integers to untagged ones going away, unpack and clean up the relatively complex first_to_remain = max(snap.idx + 1 - preserve_log_entries, 0) calculation in persistence::store_snapshot(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Laszlo Ersek	4dc2faa49a	test/raft/replication: clean up remaining index_t usage With implicit conversion of tagged integers to untagged ones going away, explicitly untag the operands / arguments of the following operations, in "test/raft/replication.hh": - assignment to raft_cluster::_seen - call to hasher_int::hash_range() Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Laszlo Ersek	3a32f3de81	test/raft/replication: take an "index_t start_idx" in create_log() raft_cluster::get_states() passes a "start_idx" to create_log(), and create_log() uses it as an "index_t" object. Match the type of "start_idx" to its name. This patch is best viewed with "git show -W". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Laszlo Ersek	08e117aeb5	test/raft/replication: untag index_t in test_case::get_first_val() In test_case::get_first_val(), the asssignment first_val = initial_snapshots[initial_leader].snap.idx; both relies on implicit conversion of the tagged integer type "index_t" to the underlying "uint64_t", and is a logic bug, as reported at <https://github.com/scylladb/scylladb/issues/20151>. For now, wean the buggy asssignment off the disappearing tagged-to-untaggged conversion. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Laszlo Ersek	6254fca7f5	test/raft/etcd_test: tag index_t and term_t for comparisons and subtractions Properly annotate index_t and term_t constants for use in BOOST_CHECK_EQUAL() and BOOST_CHECK(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Laszlo Ersek	bd4fc85bf0	test/raft/fsm_test: tag index_t and term_t for comparisons and subtractions Properly annotate index_t and term_t constants for use in BOOST_CHECK_EQUAL(), BOOST_CHECK(). Clean up the first args of read_quorum() calls -- stay in term_t space. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Laszlo Ersek	265655473e	test/raft/helpers: tighten compare_log_entries() param types The "from" and "to" parameters of compare_log_entries() are raft log indices; change them to raft::index_t, and update the callers. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 22:54:42 +02:00
Piotr Smaron	3e3858521d	codeowners: add appropriate reviewers to the frontend components	2024-08-14 22:26:35 +02:00
Piotr Smaron	1b2e88b96a	codeowners: fix codeowner names	2024-08-14 22:26:26 +02:00
Laszlo Ersek	5dcc627465	service/raft_sys_table_storage: tweak dead code In raft_sys_table_storage::store_snapshot_descriptor(), the condition preserve_log_entries > snap.idx both relies on implicit conversion of the tagged integer type "index_t" to the underlying "uint64_t", and is a logic bug, as reported at <https://github.com/scylladb/scylladb/issues/20080>. Ticket#20080 explains that this condition always evaluates to false in practice, and that the "else" branch handles all cases correctly anyway. For now, wean the buggy expression off the disappearing tagged-to-untaggged conversion. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 21:35:34 +02:00
Andrei Chekun	3407ae5d8f	[test.py] Add Junit logger for boost test Currently, boost tests aren't using Junit. Enable Junit report output and clean them from skipped test, since boost tests are executed by function name rather than filename. This allows including boost tests result to the Allure report. Related: https://github.com/scylladb/qa-tasks/issues/1665 Closes scylladb/scylladb#19925	2024-08-14 22:18:31 +03:00
Avi Kivity	6d6f93e4b5	Merge 'test/nodetool: enable running nodetool tests under test/nodetool' from Kefu Chai before this change, we assume user runs nodetool tests right under the root source directory. if user runs them under `test/nodetool`, the suppression rules are not applied. as the path is incorrect in that case. after this change, the supression rules' path is deduced from the top src directory. so we can now run the nodetool test under `test/nodetool` . --- no need to backport, this change improves developer's experience. Closes scylladb/scylladb#20119 * github.com:scylladb/scylladb: test/nodetool: deduce subpression path from top srcdir test/nodetool: deduce path from top srcdir	2024-08-14 22:10:38 +03:00
Michał Jadwiszczak	f7eb74e31f	cql3/statements/create_service_level: forbid creating SL starting with `$` Tenant names starting with `$` are reserved for internal ones. Forbid creating new service level which name starts with `$` and log a warning for existing service levels with `$` prefix. Closes scylladb/scylladb#20122	2024-08-14 21:25:31 +03:00
Kefu Chai	5ce07e5d84	build: cmake: add compiler-training target `tools/toolchain/optimized_clang.sh` builds this target for creating the profile in order to build clang optimized with this profile data. so let's be compatible with `configure.py`, and add this target to CMake building system as well. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20105	2024-08-14 21:21:33 +03:00
Ernest Zaslavsky	f5f65ead1e	Add `.clang-format`, also add CLion build folder to the `.gitignore` file Closes scylladb/scylladb#20123	2024-08-14 21:20:29 +03:00
Pavel Emelyanov	66d72e010c	distributed_loader: Lock table via global table ptr The lock_table() method needs database, ks and cf to find the table on all shards. The same can be achieved with the help of global_table_ptr thing that all the core callers already have at hand. There's a test that doesn't have global table, but it can get one. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20139	2024-08-14 20:53:21 +03:00
Pavel Emelyanov	7e3e5cfcad	sstable_directory: Simplify special-purpose local-only constructor Typically the sstable_directory is constructed out of a table object. Some code, namely tests and schema-loader, don't have table at hand and construct directory out of schema, sharder, path-to-sstables, etc. This code doesn't work with any storage options other than local ones, so there's no need (yet) to carry this argument over. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#20138	2024-08-14 20:22:50 +03:00
Avi Kivity	28d3b91cce	Merge 'test/perf/perf_sstables: use test_modes as the type of its option' from Kefu Chai before this change, we look up for the mode using the command line option as the key, but that's incorrect if the command line option does not match with any of the known names. in that case, `test_mode` just create another pair of <sstring, test_modes>, and return the second component of this pair. and the second component is not what we expect. we should have thrown an exception. in this change * the test_mode map is marked const. * the overloads for parsing / formatting the `test_modes` type are added, so that boost::program_options can parse and format it. after this change, we print more user friendly error, like ``` /scylla perf-sstable --mode index-foo error: the argument ('index-foo') for option '--mode' is invalid Try --help. ``` instead of a bunch of output which is printed as if we passes the correct option as the argument of the `--mode` option. --- it's an improvement of developer experience, hence no need to backport. Closes scylladb/scylladb#20140 * github.com:scylladb/scylladb: test/perf/perf_sstable: use switch-case when appropriate test/perf/perf_sstables: use test_modes as the type of its option	2024-08-14 20:18:22 +03:00
Piotr Smaron	31cb5b132b	codeowners: remove non contributors	2024-08-14 18:52:25 +02:00
Avi Kivity	3de4e8f91b	Merge 'cql: process LIMIT for GROUP BY select queries' from Paweł Zakrzewski This change fixes #17237, fixes #5361 and fixes #5362 by passing the limit value down the call chain in cql3. A test is also added. fixes #17237 fixes #5361 fixes #5362 The regression happened in 5.4 as we changed the way GROUP BY is processed in `432cb02` - to force aggregation when it is used. The LIMIT value was not passed to aggregations and thus we failed to adhere to it. W want to backport this fix to 5.4 and 6.0 to have continuous correct results for the test case from #17237 This patch consists of 4 commits: - fa4225ea0fac2057b7a9976f57dc06bcbd900cd4 - cql3: respect the user-defined page size in aggregate queries - a precondition for this patch to be implementable - 8fbe69e74dca16ed8832d9a90489ca47ba271d0b - cql3/select_statement: simplify the get_limit function - the `do_get_limit()` function did a lot of legwork that should not be associated with it. This change makes it trivial and makes its callers do additional checks (for unset guards, or for an aggregate query) - 162828194a2b88c22fbee335894ff045dcc943c9 - cql3: process LIMIT for GROUP BY queries - pass the limit value down the chain and make use of it. This is the actual fix to #17237 - b3dc6de6d6cda8f5c09b01463bb52f827a6a00b4 - test/cql-pytest: Add test for GROUP BY queries with LIMIT - tests Closes scylladb/scylladb#18842 * github.com:scylladb/scylladb: test/cql-pytest: Add test for GROUP BY queries with LIMIT cql3: process LIMIT for GROUP BY queries cql3/select_statement: simplify the get_limit function cql3: respect the user-defined page size in aggregate queries	2024-08-14 17:54:59 +03:00
Avi Kivity	8c257db283	Merge 'Native reverse pages over RPC' from Łukasz Paszkowski Drop half-reversed (legacy) format of query::partition_slice. The select query builds a fully reversed (native) slice for reversed queries and use it together with a reversed schema to construct query::read_command that is further propagated to the database. A cluster feature is added to support nodes that still operate on half-reversed slices. When the feature is turned off: - query::read_command is transformed (to have table schema and half-reversed slices) before sending to other nodes - query::read_command is transformed (to have query schema (reversed) and reversed slices) after receiving it from other nodes - Similarly, mutations are transformed. They are reversed before being sent to other nodes or after receiving them from other nodes. Additional manual tests were performed to test a mixed-node cluster: 1. 3-node cluster with one node upgraded: reverse read queries performed on an old node 2. 3-node cluster with one node upgraded: reverse read queries performed on a new node 3. 3-node cluster with one node upgraded and all its sstable files deleted to trigger repair: reverse read queries performed on an old node 4. 3-node cluster with one node upgraded and all its sstable files deleted to trigger repair: reverse read queries performed on a new node All reverse read queries above consists of: - single-partition reverse reads with no clustering key restrictions, with single column restrictions and multi column restrictions both with and without paging turned on - multi-partition reverse reads with range restrictions with optional partition limit and partial ordering The exact same tests were also performed on a fully upgraded cluster. Fixes https://github.com/scylladb/scylladb/issues/12557 Closes scylladb/scylladb#18864 * github.com:scylladb/scylladb: mutation_partition: drop reverse parameter in compact_for_query clustering_key_filter: unify get_ranges and get_native_ranges streamed_mutation_freezer: drop the reverse parameter reverse-reads.md: Drop legacy reverse format information Fix comments refering to half-reversed (legacy) slices select_statement::do_execute: Add tracing informaction query::trim_clustering_row_ranges_to: require reversed schema for native reversed ranges query-request: Drop half_reverse_slice as it is no longer used anywhere readers: Use reversed schema and native reversed slices database: accept reversed schema for reversed queries storage_proxy: Support reverse queries in native format query_pagers: Replace _schema with _query_schema query_pagers: Support reverse queries in native format select_statement: Execute reversed query in native format storage_proxy::remote: Add support for mixed-node clusters mutation_query: Add reversed function to reverse reconcilable_result query-request: Add reversed function to reverse read_command features: add native_reverse_queries kl::reader::make_reader: Unify interface with mx::reader::make_reader config: drop reversed_reads_auto_bypass_cache config: drop enable_optimized_reversed_reads	2024-08-14 17:51:56 +03:00
Anna Stuchlik	99be8de71e	doc: set 6.1 as the latest stable version This commit updates the configuration for ScyllaDB documentation so that: - 6.1 is the latest version. - 6.1 is removed from the list of unstable versions. It must be merged when ScyllaDB 6.1 is released. No backport is required. Closes scylladb/scylladb#20041	2024-08-14 13:43:17 +02:00
Laszlo Ersek	d87d1ae29d	service/raft_sys_table_storage: simplify (snap.idx - preserve_log_entries) With conversion of tagged integers to untagged ones going away, replace static_cast<uint64_t>(snap.idx) with snap.idx.value() Furthermore, casting "preserve_log_entries" (of type "size_t") to "uint64_t" is redundant (both "snap.idx" and "preserve_log_entries" carry nonnegative values, and the mathematical difference is expected to be nonnegative); remove the cast. Finally, simplify the initialization syntax. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	e781046739	service/raft_sys_table_storage: untag index_t and term_t for queries With implicit conversion of tagged integers to untagged ones going away, explicitly untag index_t and term_t values in the following two contexts: - when they are passed to CQL queries as int64_t, - when they are default-constructed as fallbacks for int64_t fields missing from CQL result sets. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	4f1f207be1	raft/server: clean up index_t usage With implicit conversion of tagged integers to untagged ones going away, explicitly tag (or untag, as necessary) the operands of the following operations, in "raft/server.cc": - addition of tagged and untagged (both should be tagged) - subscripting an array by tagged (should be untagged) - comparing a size-like threshold against tagged (should be untagged) - exposing tagged via gauges (should be untagged) Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	1b134d52ac	raft/tracker: don't drop out of index_t space for subtraction Tagged integers support subtraction; use it. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	b6233209d9	raft/fsm: clean up index_t and term_t usage With implicit conversion of tagged integers to untagged ones going away, explicitly tag (or untag, as necessary) the operands of the following operations, in "raft/fsm.cc": - addition of tagged and untagged (both should be tagged) - comparison (relop) between tagged an untagged (both should be tagged) - subscripting or sizing an array by tagged (should be untagged) Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	5b9a4428c6	raft/log: clean up index_t usage With implicit conversion of tagged integers to untagged ones going away, explicitly tag (or untag, as necessary) the operands of the following operations, in raft/log.{cc,h}: - addition of tagged and untagged (both should be tagged) - comparison (relop) between tagged an untagged (both should be tagged) - subscripting an array, or offsetting an iterator, by tagged (should be untagged) - comparing an array bound against tagged (should be untagged) - subtracting tagged from an array bound (should be untagged) Note: these files mix uniform initialization syntax (index_t{...}) with constructor call syntax (index_t()), with the former being more frequent. Stick with the former here too, for consistency. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	9e95f3a198	db/system_keyspace: promise a tagged integer from increment_and_get_generation() Internally, increment_and_get_generation() produces a "gms::generation_type" value. In turn, all callers of increment_and_get_generation() -- namely scylla_main() [main.cc] and single_node_cql_env::run_in_thread() [test/lib/cql_test_env.cc] -- pass the resolved value to storage_service::init_address_map() and storage_service::join_cluster(), both of which take a "gms::generation_type". Therefore it is pointless to "untag" the generation value temporarily between the producer and the consumers. Correct the return type of increment_and_get_generation(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	baccbc09c5	gms/gossiper: return "strong_ordering" from compare_endpoint_startup() The callers of gossiper::compare_endpoint_startup() need not (should not) learn of any particular (tagged or untagged) difference of generations; they only care about the ordering of generations. Change the return type of compare_endpoint_startup() to "std::strong_ordering", and delegate the comparison to tagged_tagged_integer::operator<=>. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Laszlo Ersek	3bb608056c	gms/gossiper: get "int32_t" value of "gms::version_type" explicitly In do_sort(), we need to drop to "int32_t" temporarily, so that we can call ::abs() on the version difference. Do that explicitly. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-14 13:35:08 +02:00
Michał Chojnowski	4d77faa61e	cql_test_env: ensure shutdown() before stop() for system_keyspace If system_keyspace::stop() is called before system_keyspace::shutdown(), it will never finish, because the uncleared shared pointers will keep it alive indefinitely. Currently this can happen if an exception is thrown before the construction of the shutdown() defer. This patch moves the shutdown() call to immediately before stop(). I see no reason why it should be elsewhere. Fixes scylladb/scylla-enterprise#4380 Closes scylladb/scylladb#20089	2024-08-14 12:16:44 +03:00
Kefu Chai	06c60f6abe	test/perf/perf_sstable: use switch-case when appropriate instead of using a chain of `if-else`, use switch-case instead, it's visually easier to follow than `if`-`else` blocks. and since we never need to handle the `else` case, the `throw` statement is removed. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-14 17:14:42 +08:00
Kefu Chai	5141c6efe0	test/perf/perf_sstables: use test_modes as the type of its option before this change, we look up for the mode using the command line option as the key, but that's incorrect if the command line option does not match with any of the known names. in that case, `test_mode` just create another pair of <sstring, test_modes>, and return the second component of this pair. and the second component is not what we expect. we should have thrown an exception. in this change * the test_mode map is marked const. * the overloads for parsing / formatting the `test_modes` type are added, so that boost::program_options can parse and format it. after this change, * we can print more user friendly error, like ``` /scylla perf-sstable --mode index-foo error: the argument ('index-foo') for option '--mode' is invalid Try --help. ``` instead of a bunch of output which is printed as if we passes the correct option as the argument of the `--mode` option. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-14 17:14:42 +08:00
Dawid Medrek	4ba9cb0036	README: Update the version of C++ to C++23 Scylla has started being built with C++23. We update the information in the relevant documents accordingly. Closes scylladb/scylladb#20134	2024-08-14 12:06:23 +03:00
Kamil Braun	a3d53bd224	Merge 'Prevent ALTERing non-existing KS with tablets' from Piotr Smaron ALTER tablets KS executes in 2 steps: 1. ALTER KS's cql handler forms a global topo req, and saves data required to execute this req, 2. global topo req is executed by topo coordinator, which reads data attached to the req. The KS name is among the data attached to the req. There's a time window between these steps where a to-be-altered KS could have been DROPped, which results in topo coordinator forever trying to ALTER a non-existing KS. In order to avoid it, the code has been changed to first check if a to-be-altered KS exists, and if it's not the case, it doesn't perform any schema/tablets mutations, but just removes the global topo req from the coordinator's queue. BTW. just adding this extra check resulted in broader than expected changes, which is due to the fact that the code is written badly and needs to be refactored - an effort that's already planned under #19126 (I suggest to disable displaying whitespace differences when reviewing this PR). Fixes: scylladb/scylladb#19576 Closes scylladb/scylladb#19666 * github.com:scylladb/scylladb: tests: ensure ALTER tablets KS doesn't crash if KS doesn't exist cql: refactor rf_change indentation Prevent ALTERing non-existing KS with tablets	2024-08-14 10:27:41 +02:00
Piotr Smaron	ddb5204929	tests: ensure ALTER tablets KS doesn't crash if KS doesn't exist Using the error injection framework, we inject a sleep into the processing path of ALTER tablets KS, so that the topology coordinator of the leader node sleeps after the rf_change event has been scheduled, but before it is started to be executed. During that time the second node executes a DROP KS statement, which is propagated to the leader node. Once leader node wakes up and resumes processing of ALTER tablets KS, the KS won't exist and the node cannot crash, which was the case before.	2024-08-13 21:51:51 +02:00
Pavel Emelyanov	05adee4c82	test: Add test for s3::client::bucket_lister Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 21:15:43 +03:00
Pavel Emelyanov	a02e65c649	s3_client: Add bucket lister The lister resembles the directory_lister from util -- it returns entries upon its .get() invocation, and should be .close()d at the end. Internally the lister issues ListObjectsV2 request with provided prefix and limits the server with the amount of entries returned not to consume too much local memory (we don't have streaming XML parser for response). If the result is indeed truncated, the subsequent calls include the continuation token as per [1] [1] https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjectsV2.html Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 21:15:43 +03:00
Avi Kivity	d82fd8b5f0	Merge 'Relax sstable_directory::process_descriptor() call graph' from Pavel Emelyanov The method logic is clean and simple -- load sstable from the descriptor and sort it into one of collections (local, shared, remote, unsorted). To achieve that there's a bunch of helper methods, but they duplicate functionality of each other. Squashing most of this code into process_descriptor() makes it easier to read and keeps sstable_directory private API much shorter. Closes scylladb/scylladb#20126 * github.com:scylladb/scylladb: sstable_directory: Open-code load_sstable() into process_descriptor() sstable_directory: Squash sort_sstable() with process_descriptor() sstable_directory: Remove unused sstable_filename(desc) helper sstable_directory: Log sst->get_filename(), not sstable_filename(desc) sstable_directory: Keep loaded sst in local var sstable_directory: Remove unused helpers sstable_directory: Load sstable once when sorting	2024-08-13 16:42:52 +03:00
Pavel Emelyanov	d3870304a9	sstable_directory: Open-code load_sstable() into process_descriptor() There are two load_sstable() overloads, and one of them is only used inside process_descriptor(). What this loading helper does is, in fact, processes given descriptor, so it's worth having it open-coded into its caller. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 13:27:00 +03:00
Pavel Emelyanov	da4a5df339	sstable_directory: Squash sort_sstable() with process_descriptor() The latter (caller) loads sstable, so does the former, so load it once and then put it in either list/set, depending on flags and shard info. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 13:26:10 +03:00
Pavel Emelyanov	d8cb175fb7	sstable_directory: Remove unused sstable_filename(desc) helper It's unused after previous patch. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 12:55:40 +03:00
Pavel Emelyanov	aa40aeb72f	sstable_directory: Log sst->get_filename(), not sstable_filename(desc) There are some places that log sstable Data file name via sstable descriptor. After previous patching all those loggers have sstable at hand and can use sstable::get_filename() instead. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 12:55:40 +03:00
Pavel Emelyanov	369f9111b8	sstable_directory: Keep loaded sst in local var This will make next patch shorter. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 12:55:40 +03:00
Pavel Emelyanov	ad3725fbbd	sstable_directory: Remove unused helpers After previous patch some wrappers around load_sstable() became unused. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 12:55:40 +03:00
Pavel Emelyanov	63f1969e08	sstable_directory: Load sstable once when sorting In order to decide which list to put sstable into, the sort_sstable() first calls get_shards_for_this_sstable() which loads the sstable anyway. If loaded shards contain only the current one (which is the common case) sstable is loaded again. In fact, if the sstable happens to be remote it's loaded anyway to get its open info. Fix that by loading sstable, then getting shards directly from it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 12:55:16 +03:00
Łukasz Paszkowski	ba2f037af5	mutation_partition: drop reverse parameter in compact_for_query The reverse parameter is no longer used with native reverse reads. The row ranges are provided in native reverse order together with a reversed schema, thus the reverse parameter remain false all the time and can be droped.	2024-08-13 10:07:12 +02:00
Łukasz Paszkowski	43221bbeed	clustering_key_filter: unify get_ranges and get_native_ranges When a reverse slice is provided, it is given in the native reverse format. Thus the ranges will be returned in the same order as stored in the slice. Therefore there is no need to distinguish between get_ranges and get_native_ranges. The latter one gets dropped and get_ranges returns ranges in the same order as stored in the slice.	2024-08-13 10:07:12 +02:00
Łukasz Paszkowski	8b5ec0e963	streamed_mutation_freezer: drop the reverse parameter The reverse parameter is no longer used with native reverse reads. A reversed schema is provided and thus the reverse parameter shall remain false all the time.	2024-08-13 10:07:12 +02:00
Łukasz Paszkowski	f4ca734ccb	reverse-reads.md: Drop legacy reverse format information	2024-08-13 10:07:12 +02:00
Łukasz Paszkowski	b3bf555036	Fix comments refering to half-reversed (legacy) slices	2024-08-13 10:07:12 +02:00
Łukasz Paszkowski	15a01c7111	select_statement::do_execute: Add tracing informaction Add information on table and query schema versions to tracing.	2024-08-13 10:07:12 +02:00
Łukasz Paszkowski	158b994676	query::trim_clustering_row_ranges_to: require reversed schema for native reversed ranges Simplify implementation and for clustering key ranges in native reversed format, require a reversed table schema. Trimming native reversed clustering key ranges requires a reversed schema to be passed in. Thus, the reverse flag is no longer required as it would always be set to false.	2024-08-13 10:07:10 +02:00
Łukasz Paszkowski	8d95d44027	query-request: Drop half_reverse_slice as it is no longer used anywhere	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	da95f44adc	readers: Use reversed schema and native reversed slices The reconcilable_result is built as it would be constructed for forward read queries for tables with reversed order. Mutations constructed for reversed queries are consumed forward. Drop overloaded reversed functions that reverse read_command and reconcilable_result directly and keep only those requiring smart pointers. They are not used any more.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	faa62310d9	database: accept reversed schema for reversed queries Remove schema reversing in query() and query_mutations() methods. Instead, a reversed schema shall be passed for reversed queries. Rename a schema variable from s into query_schema for readability.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	df734e35a1	storage_proxy: Support reverse queries in native format For reversed queries, query_result() method accepts a reversed table schema and read_command with a query schema version and a slice in native reversed format. Support mixed-node clusters. In such a case, the feature flag native_reverse_queries is disabled and the read_command in sent to replicas in the old regacy format (stores table schema version and a slice in the legacy reverse format). After the reconciliation, for the read+repair case, un-reversed mutations are sent to replicas, i.e. forward ones.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	d9e76a5295	query_pagers: Replace _schema with _query_schema For readability purposes. As the constructor accepts a query schema, let the varaible holding a schema be called _query_schema.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	0b2e5ff28f	query_pagers: Support reverse queries in native format For reversed queries, accept a reversed table schema and read_command with a query schema version and a slice in native reversed format.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	309ba68692	select_statement: Execute reversed query in native format Use a reversed schema and a native reversed slice when constructing a read_command and executing a reversed select statement. Such a created read_command is passed further down to query_pagers::pager and storage::proxy::query_result that transform it to the format they accept/know, i.e. lagacy.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	8c391a8ebe	storage_proxy::remote: Add support for mixed-node clusters In handle_read, detect whether a coming read_command is in the legacy reversed format or native reversed format. The result will be used to transform the read_command between format as well as to transforms the results before they are send back to the coordinator.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	fbd324b5cd	mutation_query: Add reversed function to reverse reconcilable_result The reconcilable_result is reversed by reversing mutations for all paritions it holds. Reversing is asynchronous to avoid potential stall. Use for transitions between legacy and native formats and in order to support mixed-nodes clusters.	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	b91edbacf1	query-request: Add reversed function to reverse read_command The read_command is reversed by reversing the schema version it holds and transforming a slice from the legacy reversed format to the native reversed format. Use for trasition between format and to support mixed-nodes clusters	2024-08-13 10:03:46 +02:00
Łukasz Paszkowski	9690785112	features: add native_reverse_queries Enabled when all replicas support the native_reversed command slice and return the result in reverse order in this case.	2024-08-13 10:03:42 +02:00
Łukasz Paszkowski	7b201e9165	kl::reader::make_reader: Unify interface with mx::reader::make_reader Ensure both readers have the same interfaces to avoid mistakes as both readers are used in sstable::make_reader. Less error prone.	2024-08-13 10:02:43 +02:00
Łukasz Paszkowski	b270097f1f	config: drop reversed_reads_auto_bypass_cache Reverse reads have already been with us for a while, thus this back door option to bypass in-memory data cache for reversed queries can be retired.	2024-08-13 10:02:42 +02:00
Łukasz Paszkowski	80df313f49	config: drop enable_optimized_reversed_reads Reverse reads have already been with us for a while, thus this back door option to read entire paritions forward and reversing them after can be retired.	2024-08-13 10:02:42 +02:00
Pavel Emelyanov	6675bd8a5c	s3_client: Encode query parameter value for query-string When signing AWS query one need to prepare "query string" which is a line looking like `encode(query_param)=encode(query_value)&...`. Encoded are only the query parameter names and values. It was missing in current code and so far worked because no encodable characters were used. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-13 10:59:31 +03:00
Raphael S. Carvalho	74612ad358	tablets: Fix race between repair and split Consider the following: T 0 split prepare starts 1 repair starts 2 split prepare finishes 3 repair adds unsplit sstables 4 repair ends 5 split executes If repair produces sstable after split prepare phase, the replica will not split that sstable later, as prepare phase is considered completed already. That causes split execution to fail as replicas weren't really prepared. This also can be triggered with load-and-stream which shares the same write (consumer) path. The approach to fix this is the same employed to prevent a race between split and migration. If migration happens during prepare phase, it can happen source misses the split request, but the tablet will still be split on the destination (if needed). Similarly, the repair writer becomes responsible for splitting the data if underlying table is in split mode. That's implemented in replica::table for correctness, so if node crashes, the new sstable missing split is still split before added to the set. Fixes #19378. Fixes #19416. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-08-12 17:28:51 -03:00
Raphael S. Carvalho	239344ab55	compaction: Allow "offline" sstable to be split In order to fix the race between split and repair, we must introduce the ability to split an "offline" sstable, one that wasn't added to any of the table's sstable set yet. It's not safe to split a sstable after adding it to the set, because a failure to split can result in unsplit data left in the set, causing split to fail down the road, since the coordinator thinks this replica has only split data in the set. Refs #19378. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-08-12 17:27:16 -03:00
Laszlo Ersek	607abe96e8	test/sstable: merge test_using_reusable_sst*() All lambdas passed to test_using_reusable_sst() conform to the prototype void (test_env&, sstable_ptr) All lambdas passed to test_using_reusable_sst_returning() conform to the prototype NON_VOID (test_env&, sstable_ptr) The common parameter list of both prototypes can be expressed with the concept std::invocable<test_env&, sstable_ptr> Once a "Func" template parameter (i.e., function type) satisfying this concept is taken, then "Func"'s void or non-void return type can be commonly expressed with std::invoke_result_t<Func, test_env&, sstable_ptr> In turn, test_env::do_with_async_returning<...> can be instantiated with this return type, even if it happens to be "void". ([stmt.return] specifies, "[a] return statement with an operand of type void shall be used only in a function that has a cv void return type", meaning that return func(env) will do the right thing in the body of test_env::do_with_async_returning<void>().) Merge test_using_reusable_sst() and test_using_reusable_sst_returning() into one. Preserve the function name from the former, and the test_env::do_with_async_returning<...>() call from the latter. Suggested-by: Avi Kivity <avi@scylladb.com> Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#20090	2024-08-12 17:52:01 +03:00
Kefu Chai	db4654ca49	test/nodetool: deduce subpression path from top srcdir there are chances that developer launch `pytest` right under `test/nodetool`, in that case current working directory is not the root directory of the project, so the path to suppression rules does not point to a file. to cater the needs to run the test under `test/nodetool`, let's use the path deduced from the top_srcdir. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-12 22:50:18 +08:00
Kefu Chai	c817e13d63	test/nodetool: deduce path from top srcdir add a helper to get path from top src dir, more readable this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-08-12 22:50:18 +08:00
Nikos Dragazis	90363ce802	test: Test the SSTable validation API against malformed SSTables Unit testing for the SSTable validation API happens in `sstable_validate_test`. Currently, this test checks the API against some invalid SSTables with out-of-order clustering rows and out-of-order partitions. However, both are types of content-level corruption that do not trigger `malformed_sstable_exception` errors. Extend the test to cover cases of file-level corruption as well, i.e., cases that would raise a `malformed_sstable_exception`. Construct an SSTable with an invalid checksum to trigger this. This is part of the effort to improve scrub to handle all kinds of corruption. Fixes scylladb/scylladb#19057 Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> Closes scylladb/scylladb#20096	2024-08-12 15:09:58 +03:00
Botond Dénes	fec57c83e6	Merge 'cell_locker: maybe_rehash: ignore allocation failures' from Benny Halevy `maybe_rehash` is complimentary and is not strictly require to succeed. If it fails, it will retry on the next call, but there's no reason to throw an exception that will fail its caller, since `maybe_rehash` is called as the final step after the caller has already succeeded with its action. Minor enhancement for the error path, no backport required. Closes scylladb/scylladb#19910 * github.com:scylladb/scylladb: cell_locker: maybe_rehash: reindent cell_locker: maybe_rehash: ignore allocation failures	2024-08-12 10:54:56 +03:00
Kefu Chai	0ae04ee819	build: cmake: use $<CONFIG:cfgs> when appropriate per https://cmake.org/cmake/help/latest/manual/cmake-generator-expressions.7.html#genex:CONFIG, `cfgs` can be a comma-separated list. this is supported by CMake 3.19 and up, and our minimum required CMake version is 3.27. so let's switch over from the composition of `IN_LIST` and `CONFIG` generator expressions to a single one. simpler this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20110	2024-08-11 21:28:38 +03:00
Avi Kivity	318278ff92	Merge 'tablets: reload only changed metadata' from Botond Dénes Currently, each change to tablet metadata triggers a full metadata reload from disk. This is very wasteful, especially if the metadata change affects only a single row in the `system.tablets` table. This is the case when the tablet load balancer triggers a migration, this will affect a single row in the table, but today will trigger a full reload. We expect tablet count to potentially grow to thousands and beyond and the overhead of this full reload can become significant. This PR makes tablet metadata reload partial, instead of reloading all metadata on topology or schema changes, reload only the partitions that are affected by the change. Copy the rest from the in-memory state. This is done with two passes: first the change mutations are scanned and a hint is produced. This hint is then passed down to the reload code, which will use it to only reload parts (rows/partitions) of the metadata that has actually changed. The performance difference between full reload and partial reload is quite drastic: ``` INFO 2024-07-25 05:06:27,347 [shard 0:stat] testlog - Tablet metadata reload: full 616.39ms partial 0.18ms ``` This was measured with the modified (by this PR) `perf_tablets`, which creates 100 tables, each with 2K tablets. The test was modified to change a single tablet, then do a full and partial reload respectively, measuring the time it takes for reach. Fixes: #15294 New feature, no backport needed. Closes scylladb/scylladb#15541 * github.com:scylladb/scylladb: test/perf/perf_tablets: add tablet metadata reload perf measurement test/boost/tablets_test: add test for partial tablet metadata updates db/schema_tables: pass tablet hint to update_tablet_metadata() service/storage_service: load_tablet_metadata(): add hint parameter service/migration_listener: update_tablet_metadata(): add hint parameter service/raft/group0_state_machine: provide tablet change hint on topology change service/storage_service: topology_state_load(): allow providing change hint replica/tablets: add update_tablet_metadata() replica/tablets: fix indentation replica/tablets: extract tablet_metadata builder logic replica/tablets: add get_tablet_metadata_change_hint() and update_tablet_metadata_change_hint() locator/tablets: add tablet_map::clear_tablet_transition_info() locator/tablets: make tablet_metadata cheap to copy mutation/canonical_mutation: add key()	2024-08-11 21:27:18 +03:00
Botond Dénes	2b2db510b7	test/perf/perf_tablets: add tablet metadata reload perf measurement Measure reload perf of full reload vs. partial reload, after changing a single tablet. While at it, modify the `--tablets-per-table` parameter, so that it has a default parameter which works OOTB. The previous default was both too large (causing oversized commitlog entry errors) and not a power of two.	2024-08-11 09:53:19 -04:00
Botond Dénes	65eee200b2	test/boost/tablets_test: add test for partial tablet metadata updates	2024-08-11 09:53:19 -04:00
Botond Dénes	b886ed44a7	db/schema_tables: pass tablet hint to update_tablet_metadata() Replace the has_tablet_mutations in `merge_tables_and_views()` with a hint parameter, which is calculated in the caller, from the original schema change mutations. This hint is then forwarded to the notifier's `update_tablet_metadata()` so that subscribers can refresh only the tablet partitions that changed.	2024-08-11 09:53:19 -04:00
Botond Dénes	5bff422b54	service/storage_service: load_tablet_metadata(): add hint parameter Allowing for reloading only those parts of the tablet metadata that were actually changed.	2024-08-11 09:53:19 -04:00
Botond Dénes	2cec0d8dd1	service/migration_listener: update_tablet_metadata(): add hint parameter The hint contains information related to what exactly changed, allowing listeners to do partial updates, instead of reloading all metadata on each notification.	2024-08-11 09:53:19 -04:00
Botond Dénes	ca302d9e28	service/raft/group0_state_machine: provide tablet change hint on topology change So that when reloading tablet state metadata from the disk, only the changed parts are reloaded.	2024-08-11 09:53:19 -04:00
Botond Dénes	806ec3244a	service/storage_service: topology_state_load(): allow providing change hint So that when reloading state from disk, only changed parts are reloaded instead of all. For now, only tablets have hints implemented.	2024-08-11 09:53:18 -04:00
Botond Dénes	bb1e733fe0	replica/tablets: add update_tablet_metadata() Allows updateng tablet metadata in-place, according to the provided hint, reading and updating only the parts that actually changed.	2024-08-11 09:52:37 -04:00
Botond Dénes	66292b4baa	replica/tablets: fix indentation Left broken from the previous patch.	2024-08-11 09:52:37 -04:00
Botond Dénes	aa378c458e	replica/tablets: extract tablet_metadata builder logic So it can be reused in a new method. Indentation is left broken deliberately, to make the patch easier to read.	2024-08-11 09:52:37 -04:00
Botond Dénes	f5976aa87b	replica/tablets: add get_tablet_metadata_change_hint() and update_tablet_metadata_change_hint() Extract a hint of what a tablet mutation changed. The hint can be later used to selectively reload only the changed parts from disk. Two variants are added: * get_tablet_metadata_change_hint() - extracts a hint from a list of tablet mutations * update_tablet_metadata_change_hint() - updates an existing hint based on a single mutation, allowing for incremental hint extraction	2024-08-11 09:52:37 -04:00
Botond Dénes	54ea71f8a6	locator/tablets: add tablet_map::clear_tablet_transition_info()	2024-08-11 09:52:37 -04:00
Botond Dénes	0254cfc7d3	locator/tablets: make tablet_metadata cheap to copy Keep lw_shared_ptr<tablet_map> in the tablet map and use COW semantics. To prevent accidental changes to shared tablet_map instances, all modifications to a tablet_map have to go through a new `mutate_tablet_map()` method, which implements the copy-modify-swap idiom.	2024-08-11 09:52:37 -04:00
Botond Dénes	fb0ab3c1fb	mutation/canonical_mutation: add key() Extracts the partition key without deserializing the entire mutation.	2024-08-11 09:52:37 -04:00
Calle Wilund	e18a855abe	extensions: Add exception types for IO extensions and handle in memtable write path Fixes #19960 Write path for sstables/commitlog need to handle the fact that IO extensions can generate errors, some of which should be considered retry-able, and some that should, similar to system IO errors, cause the node to go into isolate mode. One option would of course be for extensions to simply generate std::system_errors, with system_category and appropriate codes. But this is probably a bad idea, since it makes it more muddy at which level an error happened, as well as limits the expressibility of the error. This adds three distinct types (sharing base) distinguishing permission, availabilty and configuration errors. These are treated akin to EACCESS, ENOENT and EINVAL in disk error handler and memtable write loop. Tests updated to use and verify behaviour. Closes scylladb/scylladb#19961	2024-08-11 13:52:35 +03:00
Raphael S. Carvalho	75829d75ec	replica: Fix race between split compaction and migration After removal of rwlock (`53a6ec05ed`), the race was introduced because the order that compaction groups of a tablet are closed, is no longer deterministic. Some background first: Split compaction runs in main (unsplit) group, and adds sstable to left and right groups on completion. The race works as follow: 1) split compaction starts on main group of tablet X 2) tablet X reaches cleanup stage, so its compaction groups are closed in parallel 3) left or right group are closed before main (more likely when only main has flush work to do) 4) split compaction completes, and adds sstable to left and right 5) if e.g left is closed, adjusting backlog tracker will trigger an exception, and since that happens in row cache update's execute(), node crashes. The problem manifested as follow: [shard 0: gms] raft_topology - Initiating tablet cleanup of 5739b9b0-49d4-11ef-828f-770894013415:15 on 102a904a-0b15-4661-ba3f-f9085a5ad03c:0 ... [shard 0:strm] compaction - [Split keyspace1.standard1 009e2f80-49e5-11ef-85e3-7161200fb137] Splitting [/var/lib/scylla/data/keyspace1/...] ... [shard 0:strm] cache - Fatal error during cache update: std::out_of_range (Compaction state for table [0x600007772740] not found), at: ... -------- seastar::continuation<seastar::internal::promise_base_with_type<void>, row_cache::do_update(... -------- seastar::internal::do_with_state<std::tuple<row_cache::external_updater, std::function<seastar::future<void> ()> >, seastar::future<void> > -------- seastar::internal::coroutine_traits_base<void>::promise_type -------- seastar::internal::coroutine_traits_base<void>::promise_type -------- seastar::(anonymous namespace)::thread_wake_task -------- seastar::continuation<seastar::internal::promise_base_with_type<sstables::compaction_result>, seastar::async<sstables::compaction::run(... seastar::continuation<seastar::internal::promise_base_with_type<sstables::compaction_result>, seastar::future<sstables::compaction_resu... From the log above, it can be seen cache update failure happens under streaming sched group and during compaction completion, which was good evidence to the cause. Problem was reproduced locally with the help of tablet shuffling. Fixes: #19873. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Closes scylladb/scylladb#19987	2024-08-11 11:00:19 +03:00
Botond Dénes	1f4b9a5300	Merge 'compaction: drop compaction executors' possibility to bypass task manager' from Aleksandra Martyniuk If parent_info argument of compaction_manager::perform_compaction is std::nullopt, then created compaction executor isn't tracked by task manager. Currently, all compaction operations should by visible in task manager. Modify split methods to keep split executor in task manager. Get rid of the option to bypass task manager. Closes scylladb/scylladb#19995 * github.com:scylladb/scylladb: compaction: replace optional<task_info> with task_info param compaction: keep split executor in task manager	2024-08-11 10:26:43 +03:00
Botond Dénes	0bb1075a19	Merge 'tasks: fix task handler' from Aleksandra Martyniuk There are some bugs missed in task handler: - wait_for_task does not wait until virtual tasks are done, but returns the status immediately; - wait_for_task suffers from use after return; - get_status_recursively does not set the kind of task essentials. Fix the aforementioned. Closes scylladb/scylladb#19930 * github.com:scylladb/scylladb: test: add test to check that task handler is fixed tasks: fix task handler	2024-08-11 10:23:17 +03:00
Paweł Zakrzewski	9db272c949	test/cql-pytest: Add test for GROUP BY queries with LIMIT Remove xfail from all tests for #5361, as the issue is fixed. Remove xfail from test_group_by_clustering_prefix_with_limit It references #5362, but is fixed by #17237. Refs #17237	2024-08-11 09:08:44 +02:00
Paweł Zakrzewski	e7ae7f3662	cql3: process LIMIT for GROUP BY queries Currently LIMIT not passed to the query executor at all and it was just an accident that it worked for the case referenced in #17237. This change passes the limit value down the chain.	2024-08-11 09:08:43 +02:00
Paweł Zakrzewski	3838ad64b3	cql3/select_statement: simplify the get_limit function The get_limit() function performed tasks outside of its scope - for example checked if the statement was an aggregate. This change moves the onus of the check to the caller.	2024-08-11 09:08:43 +02:00
Paweł Zakrzewski	08f3219cb8	cql3: respect the user-defined page size in aggregate queries The comment in the code already states that we should use the user-defined page size if it's provided. To avoid OOM conditions we'll use the internally defined limit as the upper bound or if no page size is provided. This change lays ground work for fixing #5362 and is necessary to pass the test introduced in #19392 once it is implemented.	2024-08-11 09:08:43 +02:00
Michał Jadwiszczak	3745d0a534	gms/feature_service: allow to suppress features This patch adds `suppress_features` error injection. It allows to revoke support for some features and it can be used to simulate upgrade process in test.py. Features to suppress are passed as injection's value, separated by `;`. Example: `PARALLELIZED_AGGREGATION;UDA_NATIVE_PARALLELIZED_AGGREGATION` Fixes scylladb/scylladb#20034 Closes scylladb/scylladb#20055	2024-08-09 19:15:19 +02:00
Kefu Chai	a78f46aad7	s3/client: customize options for input_stream before this change, we use the default options for performing read on the input. and the default options is like ```c++ struct file_input_stream_options { size_t buffer_size = 8192; ///< I/O buffer size unsigned read_ahead = 0; ///< Maximum number of extra read-ahead operations }; ``` which is not able to offer good throughput when reading from disk, when we stream to S3. so, in this change, we use options which allows better throughput. Refs `061def001d` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20074	2024-08-09 11:52:30 +03:00
Dawid Medrek	e5d01d4000	db/hints: Make commitlog use commitlog IO scheduling group Before these changes, we didn't specify which I/O scheduling group commitlog instances in hinted handoff should use. In this commit, we set it explicitly to the commitlog scheduling group. The rationale for this choice is the fact we don't want to cause a bottleneck on the write path -- if hints are written too slowly, new incoming mutations (NOT hints) might be rejected due to a too high number of hints currently being written to disk; see `storage_proxy::create_write_response_handler_helper()` for more context. Fixes scylladb/scylladb#18654 Closes scylladb/scylladb#19170	2024-08-08 16:14:07 +02:00
Piotr Dulikowski	b72906518f	Merge 'service levels: update connections parameters automatically' from Michał Jadwiszczak This patch makes all cql connections update theirs service level parameters automatically when: - any service level is created or changed - one role is granted to another - any service level is attached to/detached from a role First of all, the patch defines what a service level and an effective service level are `938aa10509`. No new type of service levels are introduced, the commit only clarifies definitions and names what an effective service level is. (Effective service level is created by merging all service levels which are attached to all roles granted to the user. It represents exact values of connection's parameters.) Previously, to find an effective service level of a user, it required O(n) internal queries: O(n) queries to recursively find all granted roles (`standard_role_manager::query_granted()`) and a query for each role to get its service level (`standard_role_manager::get_attribute()`, which sums to O(n) queries). Because we want to reload SL parameters for all opened cql connections, we don't want to do O(n) queries for every connection, every time we create or change any service level/grant one role to another/attach or detach a service level to/from a role. To speed it up, the patch adds another layer of service level controller cache, which stored `role_name -> effective_service_level` mapping. This way finding a effective service level for a role is only a lookup to a map. Building the new cache requires only 2 queries: one to obtain all role hierarchy one to get all roles' service level. Fixes scylladb/scylladb#12923 Closes scylladb/scylladb#19085 * github.com:scylladb/scylladb: test/auth_cluster/test_raft_service_levels: add test for automatic connection update api/cql_server_test: add CQL server testing API transport/cql_server: subscribe to sl effective cache reloaded transport/controller: coroutinize `subscribe_server` and `unsubscribe_server` transport/cql_server: add method to update service level params on all connections generic_server: use async function in `for_each_gently()` service/qos/sl_controller: use effective service levels cache service/qos/service_level_controller: notify subscribers on effective cache reloaded service/raft/group0_state_machine: update effective service levels cache service/topology_coordinator: migrate service levels before auth service/qos/service_level_controller: effective service levels cache utils/sorting: allow to pass any container as verticies service/qos/service_level_controller: replace shard check to assert service/qos: define effective service level service/qos/qos_common: use const reference in `init_effective_names()` service/qos/service_level_controller: remove unused field auth: return map of directly granted roles test/auth/test_auth_v2_migration: create sl1 in the test	2024-08-08 15:31:04 +02:00
Anna Stuchlik	a1b4357765	doc: update Raft info in 6.1 This commit updates the Raft information regarding the Raft verification procedure. In 6.1, the procedure is no longer related to the upgrade. Fixes https://github.com/scylladb/scylladb/issues/19932 Closes scylladb/scylladb#20040	2024-08-08 11:25:50 +02:00
PeterFlockhart	0f9c6d24cf	Update SELECT grammar to define group_by_clause explicitly Closes scylladb/scylladb#20046	2024-08-08 12:23:20 +03:00
Avi Kivity	12c68bcf75	Merge 'querier: include cell stats in page stats' from Botond Dénes We have two mechanism to give visibility into reads having to process many tombstones: * a warning in the logs, triggered if a read processed more the `tombstone_warn_threshold` dead rows/tombstones * a trace message, which includes stats of the amount of rows in the page, including the amount of live and dead rows as well as tombstones This series extends this to also include information on cells, so we have visibility into the case where a read has to process an excessive amount of cell tombstones (mainly because of collections). A log line is now also logged if the amount of dead cells/tombstones in the page exceeds `tombstone_warn_threshold`. The trace message is also extended to contain cell stats. The `tombstone_warn_threshold` log lines now receive a 10s rate-limit to avoid excessive log spamming. The rate-limit is separate for the row and cell logs. Example of the new log line (`tombstone_warn_threshold=10` ): ``` WARN 2024-05-30 07:56:44,979 [shard 0:stmt] querier - Read 98 live cells and 126 dead cells/tombstones for system_schema.scylla_tables <partition-range-scan> (-inf, +inf) (see tombstone_warn_threshold) ``` Example of the new tracing message: ``` Page stats: 1 partition(s), 0 static row(s) (0 live, 0 dead), 1 clustering row(s) (1 live, 0 dead), 0 range tombstone(s) and 13 cell(s) (1 live, 12 dead) [shard 0] \| 2024-05-30 08:13:19.690803 \| 127.0.0.1 \| 6114 \| 127.0.0.1 ``` Fixes: https://github.com/scylladb/scylladb/issues/18996 Improvement, not a backport candidate. Closes scylladb/scylladb#18997 * github.com:scylladb/scylladb: test/boost: mutation_test: add test for cell compaction stats mutation/compact_and_expire_result: drop operator bool() querier: consume_page(): add rate-limiting to tombstone warnings querier: consume_page(): add cell stats to page stats trace message querier: consume_page(): add tombstone warning for cell tombstones querier: consume_page(): extract code which logs tombstone warning mutation/mutation_compactor: collect and aggregate cell compaction stats mutation: row::compact_and_expire(): use compact_and_expire_result collection_mutation: compact_and_expire(): use compact_and_expire_result mutation: introduce compact_and_expire_result	2024-08-08 12:16:13 +03:00
Calle Wilund	d6742e9bce	distributed_loader: Remove load_prio_keyspaces Fixes #13334 All required code paths (see enterprise) now uses extensions::is_extension_internal_keyspace. The old mechanism can be removed. One less global var. Closes scylladb/scylladb#20047	2024-08-08 12:10:27 +03:00
Avi Kivity	db77b5bd03	Merge 'convert the rest of `test/boost/sstable_test.cc` to co-routines and seastar::thread' from Laszlo Ersek This is a followup to #19937, for #19803. See in particular [this comment](https://github.com/scylladb/scylladb/issues/19803#issuecomment-2258371923). The primary conversion target is coroutines. However, while coroutines are the most convenient style, they are only infrequently usable in this case, for the following reasons: - Wherever we have a `future::finally()` that calls a cleanup function that returns a future (which must be awaited), we cannot use `co_await`. We can only use `seastar::async()` with `deferred_close` or `defer()`. - The code passes lots of lambdas, and `co_await` cannot be used in lambdas. First, I tried, and the compiler rejects it; second, a capturing lambda that is a coroutine is a trap [[1]](https://devblogs.microsoft.com/oldnewthing/20211103-00/?p=105870) [[2]](https://isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines#Rcoro-capture). In most cases, I didn't have to use naked `seastar::async()`; there were specialized wrappers in place already. Thus, most of the changes target `seastar::thread` context under existent `seastar::async()` wrappers, and only a few functions end up as coroutines. The last patch in the series (`test/sstable: remove useless variable from promoted_index_read()`) is an independent micro-cleanup, the opportunity for which I thought to have noticed while reading the code. The tail of `test/boost/sstable_test.cc` (the stuff following `promoted_index_read()`) is already written as `seastar::thread`. That's already better (for readability) than future chaining; but could have I perhaps further converted those functions to coroutines? My answer was "no": - Some of the candidate functions relied on deferred cleanups that might need to yield (all three variants of `count_rows()`). - Some had been implemented by passing lambdas to wrappers of `seastar::async()` (`sub_partition_read()`, `sub_partitions_read()`). - The test case `test_skipping_in_compressed_stream()` initially looked promising for co-routinization (from its starting point `seastar::async()`), because it seemed to employ no deferred cleanup (that might need to yield). However, the function uses three lambdas that must be able to yield internally, and one of those (`make_is()`) is even capturing. - The rest (`test_empty_key_view_comparison()`, `test_parse_path_good()`, `test_parse_path_bad()`) was synchronous code to begin with. ``` test/boost/sstable_test.cc \| 188 +++++++++----------- 1 file changed, 83 insertions(+), 105 deletions(-) ``` Refactoring; no backport needed. Closes scylladb/scylladb#20011 * github.com:scylladb/scylladb: test/sstable: remove useless variable from promoted_index_read() test/sstable: rewrite promoted_index_read() with async() test/sstable: unfuturize lambda invocation in test_using_reusable_sst() test/sstable: rewrite wrong_range() with async() test/sstable: simplify not_find_key_composite_bucket0() under test_using_reusable_sst() test/sstable: rewrite full_index_search() with async() test/sstable: simplify find_key(), all_in_place() under test_using_reusable_sst() test/sstable: rewrite (un)compressed_random_access_read() with async() test/sstable: simplify write_and_validate_sst() test/sstable: simplify check_toc_func() under async() test/sstable: simplify check_statistics_func() under async() test/sstable: simplify check_summary_func() under async() test/sstable: coroutinize check_component_integrity() test/sstable: rewrite write_sst_info() with async() test/sstable: simplify missing_summary_first_last_sane() test/sstable: coroutinize summary_query_fail() test/sstable: rewrite summary_query() with async() test/sstable: coroutinize (simple/composite)_index_read() test/sstable: rewrite index_read() with async() test/sstable: rewrite test_using_reusable_sst() with async() test/sstable: rewrite test_using_working_sst() with async()	2024-08-08 11:55:37 +03:00
Michał Jadwiszczak	b62a8b747a	test/auth_cluster/test_raft_service_levels: add test for automatic connection update	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	870bdaa6b1	api/cql_server_test: add CQL server testing API Add a CQL server testing API with and endpoint to dump service level parameters of all CQL connections. This endpoint will be later used to test functionality of automated updating CQL connections parameters.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	c3e8778ad4	transport/cql_server: subscribe to sl effective cache reloaded Make cql server (but not maintenance server) is subscribed to qos configuration change. Trigger update of connections' service level params on effective cache reloaded event. It's not done on maintenance server because it doesn't support role hierarchy nor attaching service levels.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	b2f2288292	transport/controller: coroutinize `subscribe_server` and `unsubscribe_server`	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	4af90726b6	transport/cql_server: add method to update service level params on all connections Trigger update of service level param on every cql connection. In enterprise, the method needs also to update connections' scheduling group.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	324b3c43c0	generic_server: use async function in `for_each_gently()` In the following patch, we will add a method to update service levels parameters for each cql connections. To support this, this patch allows to pass async function as a parameter to `for_each_gently()` method.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	93e6de0d04	service/qos/sl_controller: use effective service levels cache Use cache to quickly access effective service level of a role.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	664a1913c6	service/qos/service_level_controller: notify subscribers on effective cache reloaded Add event representing reload of effective service level cache and notify subscribers when the cache is reloaded.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	5f8132c13c	service/raft/group0_state_machine: update effective service levels cache Updates to `system.role_members` and `system.role_attributes` affect effective service levels cache, so applying mutations to those tables should reload the effective SL cache.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	7b28df9b4d	service/topology_coordinator: migrate service levels before auth Effective service level cache will be updated when mutations are applied to some of the auth tables. But the effective cache depends on first-level service levels cache, so service levels data should be migrated before auth data.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	842573d0af	service/qos/service_level_controller: effective service levels cache Add a second layer of service_level_controller cache which contains role name -> effective service level mapping. To build the mapping, controller uses first cache layer (service level name -> service level) and 2 queries to auth tables (one to `roles` and one to `role_members`).	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	4922f87fed	utils/sorting: allow to pass any container as verticies The container containing all verticies doesn't have to be a vector. Allowing to pass any container that meet conditions, will make to function more flexible.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	619937c466	service/qos/service_level_controller: replace shard check to assert The cache is only updated on shard 0, so doing assert is a better sanity check.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	be4c83ad3c	service/qos: define effective service level Write down definitions of `service level` and `effective service level` in service/qos/service_level_controller.hh. Until now, effective service level was only used as result of `LIST EFFECTIVE SERVICE LEVEL OF <role>`. Now we want to have quick access to effective service level of each role and introduce cache of effective sl to do it. New definitions clarify things. The commit also renames: - `update_service_levels_from_distributed_data` -> `update_service_levels_cache` Later we will introduce effective_service_level_cache, so this change standarizes the names. - `find_service_level` -> `find_effective_service_level` The function actualy returns effective service level.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	0da979e013	service/qos/qos_common: use const reference in `init_effective_names()` `service_level_options::init_effective_names()` method's argument has no reason to be mutable reference. This commit converts it to const ref.	2024-08-08 10:42:09 +02:00
Michał Jadwiszczak	37cd998993	service/qos/service_level_controller: remove unused field	2024-08-08 10:42:08 +02:00
Michał Jadwiszczak	f9048de0ce	auth: return map of directly granted roles Returns multimap of directly granted roles for each role. Uses only one query to create the map, instead of doing recursive queries for each individual role.	2024-08-08 10:42:08 +02:00
Michał Jadwiszczak	d643d5637c	test/auth/test_auth_v2_migration: create sl1 in the test Test `test_auth_v2_migration` creates auth data where role `users` has assigned service level `sl:fefe` but the service level isn't actually created. In following patches, we are going to introduce effective service levels cache which depends on auth and is refreshed when mutations are applied to v2 auth tables. Without this changes, this test will fail because the service level doesn't exist. Also the name `sl:fefe` is change to `sl1`.	2024-08-08 10:42:08 +02:00
Avi Kivity	3fe60560d2	Merge 'Coroutinize view_builder::start()' from Pavel Emelyanov It runs in the background and consists of two parts -- async() lambda and following .then()-s. This PR move the background running code into its own method and coroutinizes it in parts. With #19954 merged it finally looks really nice. Closes scylladb/scylladb#20058 * github.com:scylladb/scylladb: view_builder: Restore indentation after previous patches view_builder: Coroutinize inner start_in_background() calls view_builder: Coroutinize outer start_in_background() calls view_builder: Add helper method for background start	2024-08-07 19:47:32 +03:00
Kamil Braun	4181a1c53e	storage_service: raft topology: warn when `raft_topology_cmd_handler` fails due to abort Currently we print an ERROR on all exceptions in `raft_topology_cmd_handler`. This log level is too high, in some cases exceptions are expected -- like during shutdown. And it causes dtest failures. Turn exceptions from aborts into WARN level. Also improve logging by printing the command that failed. Fixes scylladb/scylladb#19754 Closes scylladb/scylladb#19935	2024-08-07 17:57:23 +02:00
Tomasz Grabiec	1a4baa5f9e	tablets: Do not allocate tablets on nodes being decommissioned If tablet-based table is created concurrently with node being decommissioned after tablets are already drained, the new table may be permanently left with replicas on the node which is no longer in the topology. That creates an immidiate availability risk because we are running with one replica down. This also violates invariants about replica placement and this state cannot be fixed by topology operations. One effect is that this will lead to load balancer failure which will inhibit progress of any topology operations: load_balancer - Replica 154b0380-1dd2-11b2-9fdd-7156aa720e1a:0 of tablet 7e03dd40-537b-11ef-9fdd-7156aa720e1a:1 not found in topology, at: ... Fixes #20032 Closes scylladb/scylladb#20053	2024-08-07 18:52:58 +03:00
Dawid Medrek	96509c4cf7	db/hints: Make sync points be created for all hosts when not specified Sync points are created, via POST HTTP requests, for a subset of nodes in the cluster. Those nodes are specified in a request's parameter `target_hosts`. When the parameter is empty, Scylla should assume the user wants to create a sync point for ALL nodes. Before these changes, sync points were created only for LIVE nodes. If a node was dead but still part of the cluster and the user requested creating a sync point leaving the parameter `target_hosts` empty, the dead node was skipped during the creation of the sync point. That was inconsistent with the guarantees the sync point API provides. In this commit, we fix that issue and add a test verifying that the changes have made the implementation compliant with the design of the sync point API -- the test only passes after this commit. Fixes scylladb/scylladb#9413 Closes scylladb/scylladb#19750	2024-08-07 13:15:20 +02:00
Pavel Emelyanov	63afbc0fcb	view_builder: Restore indentation after previous patches Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-07 14:00:01 +03:00
Pavel Emelyanov	aa1a5d3201	view_builder: Coroutinize inner start_in_background() calls One of the co_await-ed parts of this method is async() lambda. It can be coroutinized too. One thing to care is the semaphore units -- its scope should (?) terminate earlier than the whole start_in_background() so release it explicitly. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-07 14:00:01 +03:00
Pavel Emelyanov	167c6a9c5e	view_builder: Coroutinize outer start_in_background() calls The method consists of two parts -- one running in async() thread and continuations to it. This patch turns the latter chain into co_await-s. The mentioned chain is "guarded" by then_wrapped() catch of any exception, which is turned into a plain try-catch block. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-07 14:00:01 +03:00
Pavel Emelyanov	10a87f5c5b	view_builder: Add helper method for background start The view_builder::start() happens in the background. It's good to have explicit start_in_background() method and coroutinize it next. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-08-07 13:59:57 +03:00
Dawid Medrek	ec691a84a5	docs/hinted_handoff: Describe sync point HTTP API In this commit, we describe the mechanism of sync point in Hinted Handoff in the user documentation. We explain the motivation for it and how to use it, as well as list and describe all of the parameters involved in the process. Errors that may appear and experienced by the user are addressed in the article. Fixes scylladb/scylladb#18500 Closes scylladb/scylladb#19686	2024-08-07 11:12:23 +02:00
Pavel Emelyanov	2fd60b0adc	api: Move config-related endpoints from storage_service.cc The get_all_data_file_locations and get_saved_caches_location get the returned data from db::config and should be next other endpoints working with config data. refs: #2737 Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#19958	2024-08-07 10:18:29 +03:00
Piotr Dulikowski	1963619803	Merge 'Use cross shard barrier to start view builder' from Pavel Emelyanov When starting, view builder wants all shards to synchronize with each other in the middle of initialization. For that they all synchronize via shard-0's instance counter and a shared future. There's cross-shard barrier in utils/ that provides the same facility. Closes scylladb/scylladb#19954 * github.com:scylladb/scylladb: view_builder: Drop unused members view_builder: Use cross-shard barrier on start view_builder: Add cross-shard barrier to its .start() method	2024-08-07 08:54:15 +02:00
Botond Dénes	78206a3fad	test/boost: mutation_test: add test for cell compaction stats	2024-08-06 08:56:28 -04:00
Botond Dénes	259a59bd64	mutation/compact_and_expire_result: drop operator bool() Having an operator bool() on this struct is counter-intuitive, so this commit drops it and migrates any remaining users to bool is_live(). The purpose of this operator bool() was to help in incrementally replace the previous bool return type with compact_and_expire_result in the compact_and_expire() call stack. Now that this is done, it has served its purpose.	2024-08-06 08:56:28 -04:00
Botond Dénes	f638c37c4b	querier: consume_page(): add rate-limiting to tombstone warnings These warnings can be logged once per query, which could result in filling the logs with thousands of log lines. Rate-limit to once per 10sec.	2024-08-06 08:56:11 -04:00
Botond Dénes	d69b16a51e	querier: consume_page(): add cell stats to page stats trace message	2024-08-06 08:56:11 -04:00
Botond Dénes	98c599f73a	querier: consume_page(): add tombstone warning for cell tombstones Since it is really difficult to meaningfully aggregate cell tombstones with row tombstones, there is two separate warning for them.	2024-08-06 08:56:11 -04:00
Botond Dénes	fa2ee6d545	querier: consume_page(): extract code which logs tombstone warning Soon, we want to log a warning on too many cell tombstones as well. Extract the logging code to allow reuse between the row and cell tombstone warnings.	2024-08-06 08:56:11 -04:00
Botond Dénes	e403644c8b	mutation/mutation_compactor: collect and aggregate cell compaction stats row::compact_and_expire() now returns details cell stats. Collect and aggregate these, using the existing compaction_stats::row_stats structure.	2024-08-06 08:56:11 -04:00
Botond Dénes	0396db497c	mutation: row::compact_and_expire(): use compact_and_expire_result Collect, store and return stats about cells, via compact_and_expire_result.	2024-08-06 08:56:11 -04:00
Botond Dénes	2c6d4e21e6	collection_mutation: compact_and_expire(): use compact_and_expire_result Collect, store and return stats about cells, via compact_and_expire_result.	2024-08-06 08:56:11 -04:00
Botond Dénes	e773a8eee6	mutation: introduce compact_and_expire_result To hold cell stats, to be collected during row::compact_and_expire(). Users will come in the next patches.	2024-08-06 08:56:11 -04:00
Aleksandra Martyniuk	9ec8000499	test: add test to check that task handler is fixed	2024-08-06 13:15:33 +02:00
Aleksandra Martyniuk	811ca00cec	tasks: fix task handler There are some bugs missed in task handler: - wait_for_task does not wait until virtual tasks are done, but returns the status immediately; - wait_for_task suffers from use after return; - get_status_recursively does not set the kind of task essentials. Fix the aforementioned.	2024-08-06 13:15:13 +02:00
Anna Stuchlik	849856b964	doc: add post-installation configuration to the Web Installer page This commit extracts the information about the configuration the user should do right after installation (especially running scylla_setup) to a separate file. The file is included in the relevant pages, i.e., installing with packages and installing with Web Installer. In addition, the examples on the Web Installer page are updated with supported versions of ScyllaDB. Fixes https://github.com/scylladb/scylladb/issues/19908 Closes scylladb/scylladb#20035	2024-08-06 13:49:09 +03:00
Kamil Braun	f348f33667	raft topology: improve logging Add more logging for raft-based topology operations in INFO and DEBUG levels. Improve the existing logging, adding more details. Fix a FIXME in test_coordinator_queue_management (by readding a log message that was removed in the past -- probably by accident -- and properly awaiting for it to appear in test). Enable group0_state_machine logging at TRACE level in tests. These logs are relatively rare (group 0 commands are used for metadata operations) and relatively small, mostly consist of printing `system.group0_history` mutation in the applied command, for example: ``` TRACE 2024-08-02 18:47:12,238 [shard 0: gms] group0_raft_sm - apply() is called with 1 commands TRACE 2024-08-02 18:47:12,238 [shard 0: gms] group0_raft_sm - cmd: prev_state_id: optional(dd9d47c6-50ee-11ef-d77f-500b8e1edde3), new_state_id: dd9ea5c6-50ee-11ef-ae64-dfbcd08d72c3, creator_addr: 127.219.233.1, creator_id: 02679305-b9d1-41ef-866d-d69be156c981 TRACE 2024-08-02 18:47:12,238 [shard 0: gms] group0_raft_sm - cmd.history_append: {canonical_mutation: table_id 027e42f5-683a-3ed7-b404-a0100762063c schema_version c9c345e1-428f-36e0-b7d5-9af5f985021e partition_key pk{0007686973746f7279} partition_tombstone {tombstone: none}, row tombstone {range_tombstone: start={position: clustered, ckp{0010b4ba65c64b6e11ef8080808080808080}, 1}, end={position: clustered, ckp{}, 1}, {tombstone: timestamp=1722617232237511, deletion_time=1722617232}}{row {position: clustered, ckp{0010dd9ea5c650ee11efae64dfbcd08d72c3}, 0} tombstone {row_tombstone: none} marker {row_marker: 1722617232237511 0 0}, column description atomic_cell{ create system_distributed keyspace; create system_distributed_everywhere keyspace; create and update system_distributed(_everywhere) tables,ts=1722617232237511,expiry=-1,ttl=0}}} ``` note that the mutation contains a human-readable description of the command -- like "create system_distributed keyspace" above. These logs might help debugging various issues (e.g. when `apply` hangs waiting for read_apply mutex, or takes too long to apply a command). Ref: scylladb/scylladb#19105 Ref: scylladb/scylladb#19945 Closes scylladb/scylladb#19998	2024-08-06 11:50:16 +03:00
Kamil Braun	aa9d5fe3f5	Merge 'doc: add the 6.0-to-6.1 upgrade guide' from Anna Stuchlik This PR adds the 6.0-to-6.1 upgrade guide (including metrics) and removes the 5.4-to-6.0 upgrade guide. Compared 5.4-to-6.0, the the 6.0-to-6.1 guide: - Added the "Ensure Consistent Topology Changes Are Enabled" prerequisite. - Removed the "After Upgrading Every Node" section. Both Raft-based schema changes and topology updates are mandatory in 6.1 and don't require any user action after upgrading to 6.1. - Removed the "Validate Raft Setup" section. Raft was enabled in all 6.0 clusters (for schema management), so now there's no scenario that would require the user to follow the validation procedure. - Removed the references to the Enable Consistent Topology Updates page (which was in version 6.0 and is removed with this PR) across the docs. See the individual commits for more details. Fixes https://github.com/scylladb/scylladb/issues/19853 Fixes https://github.com/scylladb/scylladb/issues/19933 This PR must be backported to branch-6.1 as it is critical in version 6.1. Closes scylladb/scylladb#19983 * github.com:scylladb/scylladb: doc: remove the 5.4-to-6.0 upgrade guide doc: add the 6.0-to-6.1 upgrade guide	2024-08-06 10:23:18 +02:00
Andrei Chekun	cc428e8a36	[test.py] Increase pool size for CI Currently, the resource utilization in CI is low. Increasing the number of clusters will increase how many tests are executed simultaneously. This will decrease the time it takes to execute and improve resource utilization. Related: https://github.com/scylladb/qa-tasks/issues/1667 Closes scylladb/scylladb#19832	2024-08-06 11:20:36 +03:00
Botond Dénes	822d3b11d0	tool/scylla-nodetool: refresh: improve error-message on missing ks/tbl args The command has a singl check for the missing keyspace and/or table parameters and if the check fails, there is a combined error message. Apparently this is confusing, so split the check so that missing keyspace and missing table args have its own check and error message. Fixes: scylladb/scylladb#19984 Closes scylladb/scylladb#20005	2024-08-05 22:36:05 +03:00
Anna Stuchlik	32fa5aa938	doc: remove the 5.4-to-6.0 upgrade guide This commit removes the 5.4-to-6.0 upgrade guide and all references to it. It mainly removes references to the Enable Consistent Topology Updates page, which was added as enabling the feature was optional. In rare cases, when a reference to that page is necessary, the internal link is replaced with an external link to version 6.0. Especially the Handling Cluster Membership Change Failures page was modified for troubleshooting purposes rather than removed.	2024-08-05 20:13:48 +02:00
Kefu Chai	b1405da6ac	s3/client: use div_ceil() defined by utils/div_ceil.hh instead of reinventing the wheel, let's use the existing one. in this change, we trade the `div_ceil()` implementated in s3/client.cc for the existing one in utils/div_ceil.hh . because we are not using `std::lldiv()` anymore, the corresponding `#include <cstdlib>` is dropped. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20000	2024-08-05 15:35:18 +03:00
Kefu Chai	12a066ccdf	sstable_directory: use return_exception_ptr() when appropriate instead of using `std::rethrow_exception()`, use `coroutine::return_exception_ptr()` which is a little bit more efficient. See also `6cafd83e1c` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20001	2024-08-05 12:54:27 +03:00
Kefu Chai	0bc886d005	service: mark fmt::formatter<T>::format() as const fmt 11 enforces the constness of `format()` member function, if it is not marked with `const`, the tree fails to build with fmt 11, like: ``` /usr/include/fmt/base.h:1393:23: error: no matching member function for call to 'format' 1393 \| ctx.advance_to(cf.format(static_cast<qualified_type>(arg), ctx)); \| ~~~^~~~~~ /usr/include/fmt/base.h:1374:21: note: in instantiation of function template specialization 'fmt::detail::value<fmt::context>::format_custom_arg<service::migration_badness, fmt::formatter<service::migration_badness>>' requested here 1374 \| custom.format = format_custom_arg< \| ^ /home/kefu/dev/scylladb/service/tablet_allocator.cc:170:14: note: in instantiation of function template specialization 'fmt::format_to<fmt::basic_appender<char>, const locator::global_tablet_id &, const locator::tablet_replica &, const locator::tablet_replica &, const service::migration_badness &, 0>' requested here 170 \| fmt::format_to(ctx.out(), "{{tablet: {}, {} -> {}, badness: {}", candidate.tablet, candidate.src, \| ^ /home/kefu/dev/scylladb/service/tablet_allocator.cc:161:10: note: candidate function template not viable: 'this' argument has type 'const fmt::formatter<service::migration_badness>', but method is not marked const 161 \| auto format(const service::migration_badness& badness, FormatContext& ctx) { \| ^ ``` so, in this change, we mark these two `format()` member functions const. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20013	2024-08-05 12:53:42 +03:00
Piotr Dulikowski	a038a1fdef	Merge 'db: coroutinize do_apply_counter_update' from Michael Litvak rewrite the function as coroutine to make it easier to read and maintain, following lifetime issues we had and fixed in this function. The second commit adds a test that drops a table while there is a counter update operation ongoing in the table. The test reproduces issue https://github.com/scylladb/scylla-enterprise/issues/4475 and verifies it is fixed. Follow-up to https://github.com/scylladb/scylladb/pull/19948 Doesn't require backport because the fix to the issue was already done and backported. This is just cleanup and a test. Closes scylladb/scylladb#19982 * github.com:scylladb/scylladb: db: test counter update while table is dropped db: coroutinize do_apply_counter_update	2024-08-05 10:08:18 +02:00
Nadav Har'El	247b84715a	test/cql-pytest: reproducers for key length bugs Recently, some users have seen "Key size too large" errors in various places. Cassandra and Scylla impose a 64KB length limit on keys, and we have known about bugs in this area for a long time - and even had some translated Cassandra unit tests that cover some of them. But these tests did not cover all the corner cases and left us with partial and fragmented knowledge of this problem, spread over many test files and many issues. In this patch, we add a single test file, test/cql-pytest/test_key_length.py which attempts to rigourously explore the various bugs we have with CQL key length limits. These test aim to reproduce all known bugs in this area: * Refs #3017 - CQL layer accepts set values too large to be written to an sstable * Refs #10366 - Enforce Key-length limits during SELECT * Refs #12247 - Better error reporting for oversized keys during INSERT * Refs #16772 - Key length should be limited to exactly 65535, not less The following less interesting bug is already covered by many tests so I decided not to test it again: * Refs #7745 - Length of map keys and set items are incorrectly limited to 64K in unprepared CQL There's also a situation in materialized views and secondary indexes, where a column that was _not_ a key, now becomes a key, and a length limit needs to be enforced on it. We already have good test coverage for this (in test/cql-pytest/test_secondary_index.py and in test/cql-pytest/test_materialized_view.py), and we have an issue: * Refs #8627 - Cleanly reject updates with indexed values where value > 64k All 16 tests added here pass on Cassandra 5 except one that fails on https://issues.apache.org/jira/browse/CASSANDRA-19270, but 11 of the tests currently fail on Scylla (6 on #12247, 2 on #10366, 3 on #16772). It is possible that our decision in #16772 will not be to fix Scylla to match Cassandra but rather to declare that strict compatibility isn't needed in this case or even that Cassandra is wrong. But even then, having these tests which demonstrate the behavior of both Cassandra and Scylla will be important. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#16779	2024-08-05 10:13:49 +03:00
Tzach Livyatan	861a1cedea	Improve tombstone_compaction_interval description Closes scylladb/scylladb#19072	2024-08-05 10:10:55 +03:00
Pavel Emelyanov	f0f28cf685	docs: Extend debugging with info about exploring ELF notes When debugging coredumps some (small, but useful) information is hidden in the notes of the core ELF file. Add some words about it exists, what it includes and the thing that is always forgotten -- the way to get one Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#19962	2024-08-05 09:49:52 +03:00
Tzach Livyatan	858fd4d183	Update tracing.rst - fix table node_slow_log_time name Closes scylladb/scylladb#19893	2024-08-05 09:47:27 +03:00
Botond Dénes	76b6e8c5aa	Merge 'Drop datadir from keyspace::config' from Pavel Emelyanov Commit `ad0e6b79` (replica: Remove all_datadir from keyspace config) removed all_datadirs from keyspace config, now it's datadir turn. After this change keyspace no longer references any on-disk directories, only the sstables's storage driver attached to keyspace's tables does. refs #12707 Closes scylladb/scylladb#19866 * github.com:scylladb/scylladb: replica: Remove keyspace::config::datadir sstables/storage: Evaluate path for keyspace directory in storage sstables/storage: Add sstables_manager arg to init_keyspace_storage()	2024-08-05 09:46:29 +03:00
Avi Kivity	2eff4b41ad	repair: row_level: coroutinize working_row_hashes() It uses do_with, so it allocates unconditionally. Might as well use the allocation for a nice coroutine. Closes scylladb/scylladb#19915	2024-08-05 08:55:34 +03:00
Anna Stuchlik	eca2dfd8c3	doc: add OS support for version 6.1 This commit adds OS support for version 6.1 and removes OS support for 5.4 (according to our support policy for versions). Closes scylladb/scylladb#19992	2024-08-05 08:25:16 +03:00
Avi Kivity	aa1270a00c	treewide: change assert() to SCYLLA_ASSERT() assert() is traditionally disabled in release builds, but not in scylladb. This hasn't caused problems so far, but the latest abseil release includes a commit [1] that causes a 1000 insn/op regression when NDEBUG is not defined. Clearly, we must move towards a build system where NDEBUG is defined in release builds. But we can't just define it blindly without vetting all the assert() calls, as some were written with the expectation that they are enabled in release mode. To solve the conundrum, change all assert() calls to a new SCYLLA_ASSERT() macro in utils/assert.hh. This macro is always defined and is not conditional on NDEBUG, so we can later (after vetting Seastar) enable NDEBUG in release mode. [1] `66ef711d68` Closes scylladb/scylladb#20006	2024-08-05 08:23:35 +03:00
Avi Kivity	cdee667170	alternator: destroy streamed json values gently Large json return values are streamed to avoid memory pressure and stalls, but are destroyed all at once. This in itself can cause stalls [1]. Destroy them gently to avoid the stalls. [1] ++[0#1/1 100%] addr=0x46880df total=514498 count=7004 avg=73: \| seastar::backtrace<seastar::backtrace_buffer::append_backtrace_oneline()::{lambda(seastar::frame)#1}> at ./build/release/seastar.lto/./seastar/include/seastar/util/backtrace.hh:64 ++ - addr=0x4680b35: \| seastar::backtrace_buffer::append_backtrace_oneline at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:839 \| (inlined by) seastar::print_with_backtrace at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:858 ++ - addr=0x46800f7: \| seastar::internal::cpu_stall_detector::generate_trace at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:1469 ++ - addr=0x4680178: \| seastar::internal::cpu_stall_detector::maybe_report at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:1206 \| (inlined by) seastar::internal::cpu_stall_detector::on_signal at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:1226 ++ - addr=0x3dbaf: ?? ??:0 ++[1#1/812 13%] addr=0x217b774 total=69336 count=990 avg=70: \| rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericValue at /usr/include/rapidjson/document.h:721 \| ++[2#1/3 85%] addr=0x217b7db total=58974 count=842 avg=70: \| \| rapidjson::GenericMember<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericMember at /usr/include/rapidjson/document.h:71 \| \| (inlined by) rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericValue at /usr/include/rapidjson/document.h:733 \| \| ++[3#1/4 45%] addr=0x217b7db total=902102 count=12903 avg=70: \| \| \| rapidjson::GenericMember<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericMember at /usr/include/rapidjson/document.h:71 \| \| \| (inlined by) rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericValue at /usr/include/rapidjson/document.h:733 \| \| -> continued at addr=0x217b7db above \| \| \|+[3#2/4 40%] addr=0x217b8b3 total=794219 count=11363 avg=70: \| \| \| rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericValue at /usr/include/rapidjson/document.h:726 \| \| \| ++[4#1/1 100%] addr=0x217b7db total=909571 count=13012 avg=70: \| \| \| \| rapidjson::GenericMember<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericMember at /usr/include/rapidjson/document.h:71 \| \| \| \| (inlined by) rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>::~GenericValue at /usr/include/rapidjson/document.h:733 \| \| \| -> continued at addr=0x217b7db above \| \| \|+[3#3/4 15%] addr=0x43d35a3 total=296768 count=4246 avg=70: \| \| \| seastar::shared_ptr_count_for<rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator> >::~shared_ptr_count_for at ././seastar/include/seastar/core/shared_ptr.hh:492 \| \| \| (inlined by) seastar::shared_ptr_count_for<rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator> >::~shared_ptr_count_for at ././seastar/include/seastar/core/shared_ptr.hh:492 \| \| \| ++[4#1/2 98%] addr=0x43e7d06 total=289680 count=4144 avg=70: \| \| \| \| seastar::shared_ptr<rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator> >::~shared_ptr at ././seastar/include/seastar/core/shared_ptr.hh:570 \| \| \| \| (inlined by) alternator::make_streamed(rapidjson::GenericValue<rapidjson::UTF8<char>, rjson::internal::throwing_allocator>&&)::$_0::operator() at ./alternator/executor.cc:127 \| \| \| ++ - addr=0x184e0a6: \| \| \| \| std::__n4861::coroutine_handle<seastar::internal::coroutine_traits_base<void>::promise_type>::resume at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/coroutine:240 \| \| \| \| (inlined by) seastar::internal::coroutine_traits_base<void>::promise_type::run_and_dispose at ./build/release/seastar.lto/./seastar/include/seastar/core/coroutine.hh:125 \| \| \| \| (inlined by) seastar::reactor::run_tasks at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:2651 \| \| \| \| (inlined by) seastar::reactor::run_some_tasks at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:3114 \| \| \| \| ++[5#1/1 100%] addr=0x2503b87 total=310677 count=4417 avg=70: \| \| \| \| \| seastar::reactor::do_run at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:3283 \| \| \| \| ++[6#1/2 78%] addr=0x46a2898 total=400571 count=5450 avg=73: \| \| \| \| \| seastar::smp::configure(seastar::smp_options const&, seastar::reactor_options const&)::$_0::operator() at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:4501 \| \| \| \| \| (inlined by) std::__invoke_impl<void, seastar::smp::configure(seastar::smp_options const&, seastar::reactor_options const&)::$_0&> at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/invoke.h:61 \| \| \| \| \| (inlined by) std::__invoke_r<void, seastar::smp::configure(seastar::smp_options const&, seastar::reactor_options const&)::$_0&> at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/invoke.h:111 \| \| \| \| \| (inlined by) std::_Function_handler<void (), seastar::smp::configure(seastar::smp_options const&, seastar::reactor_options const&)::$_0>::_M_invoke at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/std_function.h:290 \| \| \| \| ++ - addr=0x4673fda: \| \| \| \| \| std::function<void ()>::operator() at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/std_function.h:591 \| \| \| \| \| (inlined by) seastar::posix_thread::start_routine at ./build/release/seastar.lto/./seastar/src/core/posix.cc:90 \| \| \| \| ++ - addr=0x8c946: ?? ??:0 \| \| \| \| ++ - addr=0x11296f: ?? ??:0 \| \| \| \| ++[6#2/2 22%] addr=0x2502c1e total=113613 count=1549 avg=73: \| \| \| \| \| seastar::reactor::run at ./build/release/seastar.lto/./seastar/src/core/reactor.cc:3166 \| \| \| \| ++ - addr=0x22068e0: \| \| \| \| \| seastar::app_template::run_deprecated at ./build/release/seastar.lto/./seastar/src/core/app-template.cc:276 \| \| \| \| ++ - addr=0x220630b: \| \| \| \| \| seastar::app_template::run at ./build/release/seastar.lto/./seastar/src/core/app-template.cc:167 \| \| \| \| ++ - addr=0x22334bc: \| \| \| \| \| scylla_main at ./main.cc:672 \| \| \| \| ++ - addr=0x20411cc: \| \| \| \| \| std::function<int (int, char**)>::operator() at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/std_function.h:591 \| \| \| \| \| (inlined by) main at ./main.cc:2072 \| \| \| \| ++ - addr=0x27b89: ?? ??:0 \| \| \| \| ++ - addr=0x27c4a: ?? ??:0 \| \| \| \| ++ - addr=0x28c8fb4: \| \| \| \| \| _start at ??:? Closes scylladb/scylladb#19968	2024-08-05 00:35:52 +03:00
Botond Dénes	c34127092d	reader_concurrency_semaphore: test constructor: don't ignore metrics param The for_tests constructor has a metrics parameter defaulted to register_metrics::no, but when delegating to the other constructor, a hard-coded register_metrics::no is passed. This makes no difference currently, because all callers use the default and the hard-coded value corresponds to it. Let's fix it nevertheless to avoid any future surprises. Closes scylladb/scylladb#20007	2024-08-04 21:14:42 +03:00
Laszlo Ersek	0933a52c0b	test/sstable: remove useless variable from promoted_index_read() The large_partition_schema() call returns a copy of the "schema_ptr" object that points to an effectively statically initialized thread_local "schema" object. The large_partition_schema() call has no bearing on whether, or when, the "schema" object is constructed, and has no side effects (other than copying an "lw_shared_ptr" object). Furthermore, the return value of large_partition_schema() is not used for anything in promoted_index_read(). This redundant call seems to date back to original commit `3dd079fb7a` ("tests: add test for reading parts of a large partition", 2016-08-07). Remove the call and the variable. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	bb58446258	test/sstable: rewrite promoted_index_read() with async() For better readability, replace future::then() chaining with future::get(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	1f565626d4	test/sstable: unfuturize lambda invocation in test_using_reusable_sst() All lambdas passed to test_using_reusable_sst() and test_using_reusable_sst_returning() have been converted to future::get() calls (according to the seastar::thread context that they are now executed in). None of the lambdas return futures anymore; they all directly return void or non-void. Therefore, drop futurize_invoke(...).get() around the lambda invocations in test_using_reusable_sst(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	8ea881ae04	test/sstable: rewrite wrong_range() with async() For better readability, replace the future::then() chaining (and the associated manual fiddling with object lifecycles) with future::get() (and rely on seastar::thread's stack). We're already in seastar::thread context. Similarly, replace the future::finally() underlying with_closeable() with deferred_close(); with the assumption that mutation_reader::close() never fails (and is therefore safe to call in the "deferred_close" destructor). This is actually guaranteed, as mutation_reader::close() is marked "noexcept". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	e7e9a0a696	test/sstable: simplify not_find_key_composite_bucket0() under test_using_reusable_sst() According to early patch "test/sstable: rewrite test_using_reusable_sst() with async" in this series, lambdas passed to test_using_reusable_sst() are invoked: (a) less importantly here, in seastar::thread context, (b) more importantly here, futurized (temporarily so). The test case not_find_key_composite_bucket0() doesn't chain futures; therefore it needs no conversion to future::get() for purpose (a); however, we can eliminate its empty future return. Fact (b) will cover for that, until all such lambdas are converted to direct "void" returns (at which point we can remove the futurization from test_using_reusable_sst()). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	95cf16708d	test/sstable: rewrite full_index_search() with async() For better readability, replace future::then() chaining with future::get(). (We're already in seastar::thread context.) This patch is best viewed with "git show -b". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	2a27d5b344	test/sstable: simplify find_key*(), all_in_place() under test_using_reusable_sst() According to early patch "test/sstable: rewrite test_using_reusable_sst() with async" in this series, lambdas passed to test_using_reusable_sst() are invoked: (a) less importantly here, in seastar::thread context, (b) more importantly here, futurized (temporarily so). The test cases find_key_map(), find_key_set(), find_key_list(), find_key_composite(), all_in_place() don't chain futures; therefore they need no conversion to future::get() for purpose (a); however, we can eliminate their empty future returns. Fact (b) will cover for that, until all such lambdas are converted to direct "void" returns (at which point we can remove the futurization from test_using_reusable_sst()). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	d22bd93abb	test/sstable: rewrite (un)compressed_random_access_read() with async() For better readability, replace future::then() chaining with future::get(). (We're already in seastar::thread context.) Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	6e35e584c8	test/sstable: simplify write_and_validate_sst() All three lambdas passed to write_and_validate_sst() now use future::get() rather than future::then() chaining; in other words, the future::get() calls inside all these seastar::thread contexts have been pushed down to the lambdas. Change all these lambdas' return types from future<> to void. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	8819b3f134	test/sstable: simplify check_toc_func() under async() The lambda passed to write_and_validate_sst() already runs in seastar::thread context; replace future::then() chaining with future::get() calls. We're going to eliminate the trailing "return make_ready_future<>()" later. This patch is best viewed with "git show -W -b". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	de56883a17	test/sstable: simplify check_statistics_func() under async() The lambda passed to write_and_validate_sst() already runs in seastar::thread context; replace future::then() chaining with future::get() calls. We're going to eliminate the trailing "return make_ready_future<>()" later. This patch is best viewed with "git show -W -b". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	1a85412f96	test/sstable: simplify check_summary_func() under async() The lambda passed to write_and_validate_sst() already runs in seastar::thread context; replace future::then() chaining with future::get() calls. We're going to eliminate the trailing "return make_ready_future<>()" later. This patch is best viewed with "git show -W -b". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	7b21bce1ca	test/sstable: coroutinize check_component_integrity() check_component_integrity() does not rely on any deferred close or stop operations; turn it into a coroutine therefore, for best readability. This conversion demonstrates particularly well how much the stack eases coding. We no longer need to artificially extend the lifetime of "tmp" with a final .then([tmp] {}) future. Consequently, "tmp" no longer needs to be a shared pointer to an on-heap "tmpdir" object; "tmp" can just be a "tmpdir" object on the stack. While at it, eliminate the single-use local objects "s" and "gen", for movability's sake. (We could use std::move() on these variables, but it seems easier to just flatten the function calls that produce the corresponding rvalues into the write_sst_info() argument list.) Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	caca13fe28	test/sstable: rewrite write_sst_info() with async() For better readability, replace future::then() chaining with future::get(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	cfe92ee203	test/sstable: simplify missing_summary_first_last_sane() The lambda passed to test_using_reusable_sst() is now invoked -- futurized, transitorily -- in seastar::thread context; stop returning an explicit make_ready_future<>() from the lambda. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	10ebc0a2d2	test/sstable: coroutinize summary_query_fail() summary_query_fail() does not rely on any deferred close or stop operations; turn it into a coroutine therefore, for best readability. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	a403ad0703	test/sstable: rewrite summary_query() with async() For better readability, replace future::then() chaining with future::get(). (We're already in seastar::thread context.) Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	3a57a7cfea	test/sstable: coroutinize (simple/composite)_index_read() simple_index_read() and composite_index_read() do not rely on any deferred close or stop operations; turn them into coroutines therefore, for best readability. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	eeeab1110a	test/sstable: rewrite index_read() with async() For better readability, replace future::then() chaining with future::get(). (We're already in seastar::thread context.) Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	17d4fac669	test/sstable: rewrite test_using_reusable_sst() with async() Improve the readability of test_using_reusable_sst() by replacing future::then() chaining with test_env::do_with_async() and future::get(). Unlike seastar::async(), test_env::do_with_async() restricts its input lambda to returning "void". Because of this, introduce the variant test_using_reusable_sst_returning(), based on test_env::do_with_async_returning(), for lambdas returning non-void. Put the latter to use in index_read() at once. Subsequently, we'll gradually convert the lambdas passed to test_using_reusable_sst() and test_using_reusable_sst_returning() from returning futures to returning direct values. In order for test_using_reusable_sst() and test_using_reusable_sst_returning() to cope with both types of lambdas, wrap the lambdas into futurize_invoke().get(). In the seastar::thread context, future::get() will gracefully block on genuine futures, and return immediately on direct values that were futurized on the spot. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Laszlo Ersek	79a8a6c638	test/sstable: rewrite test_using_working_sst() with async() Make test_using_working_sst() easier to read by: (1) replacing test_env::do_with() with seastar::async(), seastar::defer(), and future::get(); (2) replacing seastar::async() and seastar::defer() with test_env::do_with_async(). Technically speaking, this change does not perfectly preserve exceptional behavior. Namely, test_env::do_with() uses future::finally() to link test_env::stop() to the chain of futures, and future::finally() permits test_env::stop() itself to throw an exception -- potentially leading to a seastar::nested_exception being thrown, which would carry both the original exception and the one thrown by test_env::stop(). Contrarily, the test_env::stop() deferred with seastar::defer() runs in a destructor, and therefore test_env::stop() had better not throw there. However, we will assume that test_env::stop() does not throw, albeit not marked "noexcept". Prior commits `8d704f2532` ("sstable_test_env: Coroutinize and move to .cc test_env::stop()", 2023-10-31) and `2c78b46c78` ("sstables::test_env: Carry compaction manager on board", 2023-10-31) show that we've considered individual actions in test_env::stop() not to throw before. The 128KB stack of seastar::thread (which underlies seastar::async()) should be a tolerable cost in a test case, in exchange for the improved readability. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-08-04 15:35:51 +02:00
Kefu Chai	0660675387	utils/div_ceil: add constraints to template arguments to better reflect what we expect from the arguments. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#20003	2024-08-04 15:32:01 +03:00
Aleksandra Martyniuk	2ab56b7f56	repair: use find_column_family in insert_repair_meta repair_service::insert_repair_meta gets the reference to a table and passes it to continuations. If the table is dropped in the meantime, the reference becomes invalid. Use find_column_family at each table occurrence in insert_repair_meta instead. Closes scylladb/scylladb#19953	2024-08-04 13:56:38 +03:00
Kefu Chai	571ae0ac96	docs: link to current document instead of the github wiki before this change, the hyper link brings us to a GitHub wiki page, which just points the reader to https://docs.scylladb.com/operating-scylla/snitch/. this is not a great user experience. so, in this change, we just reference the document in the current build. more efficient this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19952	2024-08-04 11:47:21 +03:00
Kefu Chai	f7556edc65	build: cmake: define SCYLLA_ENABLE_PREEMPTION_SOURCE for dev build in `fabab2f4`, we introduced preemption_source, and added `SCYLLA_ENABLE_PREEMPTION_SOURCE` preprocessor macro to enable opt-in the pluggable preemption check. but CMake building system was not updated accordingly. so, in this change, let's sync the CMake building system with `configure.py`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19951	2024-08-04 11:46:28 +03:00
Yaron Kaikov	8221a178d8	Revert "dist: support nonroot and offline mode for scylla-housekeeping" This reverts commit `c3bea539b6`. Since it breaking offline-installer artifact-tests. Also, it seems that we should have merged it in the first place since we don't need scylla-housekeeping checks for offline-installer Closes scylladb/scylladb#19976	2024-08-04 10:55:26 +03:00
Aleksandra Martyniuk	c456a43173	compaction: replace optional<task_info> with task_info param compaction_manager::perform_compaction does not create task manager task for compaction if parent_info is set to std::nullopt. Currently, we always want to create task manager task for compaction. Remove optional from task info parameters which start compaction. Track all compactions with task manager.	2024-08-02 14:38:46 +02:00
Aleksandra Martyniuk	108d0344b8	compaction: keep split executor in task manager If perform_compaction gets std::nullopt as a parent info then the executor won't be tracked by task manager. Modify storage_group::split call so that it passes empty task_info instead of nullopt to track split.	2024-08-02 12:45:32 +02:00
Wojciech Mitros	543dab9e88	mv: test the view update behavior With the recently added mv admission control, we can now test how are the view update backlogs updated and propagated without relying just on the response delays that it was causing until now. This patch adds a test for it, replicating issues scylladb/scylladb#18461 and scylladb/scylladb#18783. In the test, we start with an empty view update backlog, then perform a write to it, increasing its backlog and saving the updated backlog on coordinator, the backlog then drops back to 0, we wait 1s for the backlog to be gossiped and we perform another write which should succeed. Due to scylladb/scylladb#18461, the test would fail because in both gossip rounds before and after the write, the backlog was empty, causing the write to be blocked by admission control indefinitely. Due to scylladb/scylladb#18783, the test would fail because when the backlog drops back to 0 after the write, the change is never registered, causing all writes to be blocked as well.	2024-08-02 12:12:24 +02:00
Wojciech Mitros	795ac177c2	mv: add test for admission control In this patch we add 2 tests for checking that the mv admission control works. The first one simply checks whether, after increasing the backlog on one node over the admission control threshold, the following request is rejected with the error message corresponding to the admission control. The second one checks whether, after triggering admission control, the entire user request fails instead of just failing a replica write. This is done by performing a number of writes, some of which trigger the admission control and cause retries, then checking if the node that had a large view update backlog received all the writes. Before, the writes would succeed on enough replicas, reaching QUORUM, and allowing the user write to succeed and cause no retries, even though on the replica with a high backlog the write got rejected due to the backlog size.	2024-08-02 12:12:24 +02:00
Wojciech Mitros	a55b7688b6	storage_proxy: return overloaded_exception instead of throwing To avoid an expensive stack unwind, instead of throwing an error, we can just return it thanks to the boost::result type that the affected methods use. The result with an exception needs to be constructed not implicitly, but with boost::outcome_v2::failure, because the exception, converted into coordinator_exception_container can be then converted into both into a successful response_id_type as well as into a failure.	2024-08-02 12:12:24 +02:00
Wojciech Mitros	5eaae05aaf	mv: reject user requests by coordinator when a replica is overloaded by MVs Currently, when a replica's view update backlog is full, the write is still sent by the coordinator to all replicas. Because of the backlog, the write fails on the replica, causing inconsistency that needs to be fixed by repair. To avoid these inconsistencies, this patch adds a check on the coordinator for overloaded replicas. As a result, a write may be rejected before being sent to any replicas and later retried by the user, when the replica is no longer overloaded. Fixes scylladb/scylladb#17426	2024-08-02 12:12:19 +02:00
Piotr Dulikowski	39b49a41cc	Merge 'mv: delete a partition in a single operation when applicable' from Michael Litvak Currently when a partition is deleted from the base table, we generate a row tombstone update for each one of the view rows in the partition. When the partition key in the view is the same as the base, maybe in a different order, this can be done more efficiently - The whole corresponding view partition can be deleted with one partition tombstone update. With this commit, when generating view updates, if the update mutation has a partition tombstone then for the views which have the same partition key we will generate a partition tombstone update, and skip the individual row tombstone updates. Fixes scylladb/scylladb#8199 Closes scylladb/scylladb#19338 * github.com:scylladb/scylladb: mv: skip reading rows when generating partition tombstone update mv: delete a partition in a single operation when applicable cql-pytest: move ScyllaMetrics to util file to allow reuse	2024-08-02 11:00:18 +02:00
Michael Litvak	0f5e8c52ad	db: test counter update while table is dropped Add a test that drops a table while there is a counter update operation ongoing in the table. The test reproduces issue scylladb/scylla-enterprise#4475 and verifies it is fixed.	2024-08-01 22:23:17 +03:00
Avi Kivity	99d0aaa7d2	Merge 'tablets: load_balancer: Improve per-table balance' from Tomasz Grabiec Tablet load balancer tries to equalize tablet load between shards by moving tablets. Currently, the tablet load balancer assumes that each tablet has the same hotness. This may not be true, and some tables may be hotter than others. If some nodes end up getting more tablets of the hot table, we can end up with request load imbalance and reduced performance. In `79d0711c7e` we implemented a mitigation for the problem by randomly choosing the table whose tablet replica should be moved. This should improve fairness of movement. However, this proved to not be enough to get a good distribution of tablets. This change improves candidate selection to not relay on randomness but rather evaluating candidates with respect to the impact on load imbalance. Also, if there is no good candidate, we consider picking other source shards, not the most-loaded one. This is helpful because when finishing node drain we get just a few candidates per shard, all of which may belong to a single table, and the destination may already be overloaded with that table. Another shard may contain tablets of another table which is not yet overloaded on the destination. And shards may be of similar load, so it doesn't matter much which shard we choose to unload. We also consider other destinations, not the least-loaded one. This helps when draining nodes and the source node has few shard candidates. Shards on the destination may have similar load so there is more than one good destinatin candidate. By limiting ourselves to a single shard, we increase the chance that we're overload the table on that shard. The algorithm was evaluated using "scylla perf-load-balancing", which simulates a sequeunce of 8 node bootstraps and decommissions for different node and shard counts, RF, and tablet counts. For example, for the following parameters: params: {iterations=8, nodes=5, tablets1=128 (2.4/sh), tablets2=512 (9.6/sh), rf1=3, rf2=3, shards=32} The results are: Before: Overcommit (old) : init : {table1={shard=1.25 (best=1.25), node=1.00}, table2={shard=1.04 (best=1.04), node=1.00}} Overcommit (old) : worst: {table1={shard=4.00 (best=1.25), node=1.81}, table2={shard=1.25 (best=1.04), node=1.11}} Overcommit (old) : last : {table1={shard=2.50 (best=1.25), node=1.41}, table2={shard=1.25 (best=1.04), node=1.05}} After: Overcommit : init : {table1={shard=1.25 (best=1.25), node=1.00}, table2={shard=1.04 (best=1.04), node=1.00}} Overcommit : worst: {table1={shard=1.50 (best=1.25), node=1.02}, table2={shard=1.12 (best=1.04), node=1.01}} Overcommit : last : {table1={shard=1.25 (best=1.25), node=1.00}, table2={shard=1.04 (best=1.04), node=1.00}} So worst shard overcommit for table1 was reduced from 4 to 1.5. Overcommit of 4 means that the most-loaded shard has 4 times more tablets than the average per-shard load in the cluster. Also, node overcommit for table1 was reduced from 1.81 to 1.02. The magnitude of improvement depends greatly on test configurtion, so on topology and tablet distribution. The algorithm is not perfect, it finds a local optimum. In the above test, overcommit of 1.5 is not the best possible (1.25). One of the reason why the current algorithm doesn't achieve best distribution is that it works with a single movement at a time and replication constraints limit the choice of destinations. Viable destinations for remaining candidates may by only on nodes which are not least-loaded, and we won't be able to fill the least loaded node. Doing so would require more complex movement involving moving a tablet from one of the destination nodes which doesn't have a replica on the least loaded node and then replacing it with the candidate from the source node. Another limitation is that the algorithm can only fix balance by moving tablets away from most loaded nodes, and it does so due to imbalance between nodes. So it cannot fix the imbalance which is already present on the nodes if there is not much to move due to similar load between nodes. It is designed to not make the imbalance worse, so it works good if we started in a good shape. Fixes https://github.com/scylladb/scylladb/issues/16824 Closes scylladb/scylladb#19779 * github.com:scylladb/scylladb: test: perf: tablet_load_balancing: Test with higher shard and tablet counts tablets: load_balancer: Avoid quadratic complexity when finding best candidate tablets: load_balancer: Maintain load sketch properly during intra-node migration tablets: load_balancer: Use "drained" flag test: perf: tablet_load_balancing: Report load balancer stats tablets: load_balancer: Move load_balancer_stats_manager to header file tablets: load_balancer: Split evaluate_candidate() into src and dst part tablets: load_balancer: Optimize evaluate_candidate() tablets: load_balancer: Add more statistics tablets: load_balancer: Track load per table on cluster level tablets: load_balancer: Track load per table on node level tablets: load_balancer: Use a single load sketch for tracking all nodes locator: load_sketch: Introduce populate_dc() tablets: load_balancer: Modify target load sketch only when emitting migration locator: load_sketch: Introduce get_most_loaded_shard() locator: load_sketch: Introduce get_least_loaded_shard() locator: load_sketch: Optimize pick()/unload() locator: load_sketch: Introduce load_type test: perf: tablet_load_balancing: Report total tablet counts test: perf: tablet_load_balancing: Print run parameters in the single simulation case too test: perf: tablet_load_balancing: Report time it took to schedule migrations tablets: load_balancer: Log table load stats after each migration tablets: load_balancer: Log per-shard load distribution in debug level tablets: load_balancer: Improve per-table balance tablets: load_balancer: Extract check_convergence() tablets: load_balancer: Extract nodes_by_load_cmp tablets: load_balancer: Maintain tablet count per table tablets: load_balancer: Reuse src_node_info test: perf: tablet_load_balancing: Print warnings about bad overcommit test: perf: tablet_load_balancing: Allow running a single simulation test: perf: tablet_load_balancing: Report best possible shard overcommit test: perf: tablet_load_balancing: Report global shard overcommit	2024-08-01 21:12:14 +03:00
Michael Litvak	22b282f5c5	db: coroutinize do_apply_counter_update rewrite the function as coroutine to make it easier to read and maintain, following lifetime issues we had and fixed in this function.	2024-08-01 19:09:04 +03:00
Anna Stuchlik	9972e50134	doc: add the 6.0-to-6.1 upgrade guide This commit adds the 6.0-to-6.1 upgrade guide. Compared to the previous upgrade guide: - Added the "Ensure Consistent Topology Changes Are Enabled" prerequisite. - Removed the "After Upgrading Every Node" section. Both Raft-based schema changes and topology updates are mandatory in 6.1 and don't require any user action after upgrading to 6.1. - Removed the "Validate Raft Setup" section. Raft was enabled in all 6.0 clusters (for schema management), so now there's no scenario that would require the user to follow the validation procedure.	2024-08-01 14:58:14 +02:00
Piotr Smaron	0ea2128140	cql: refactor rf_change indentation	2024-08-01 14:37:53 +02:00
Piotr Smaron	5b089d8e10	Prevent ALTERing non-existing KS with tablets ALTER tablets KS executes in 2 steps: 1. ALTER KS's cql handler forms a global topo req, and saves data required to execute this req, 2. global topo req is executed by topo coordinator, which reads data attached to the req. The KS name is among the data attached to the req. There's a time window between these steps where a to-be-altered KS could have been DROPped, which results in topo coordinator forever trying to ALTER a non-existing KS. In order to avoid it, the code has been changed to first check if a to-be-altered KS exists, and if it's not the case, it doesn't perform any schema/tablets mutations, but just removes the global topo req from the coordinator's queue. BTW. just adding this extra check resulted in broader than expected changes, which is due to the fact that the code is written badly and needs to be refactored - an effort that's already planned under #19126 Fixes: #19576	2024-08-01 14:37:53 +02:00
Piotr Dulikowski	44f327675d	Merge 'Remove gossiper argument from storage_service::join_cluster()' from Pavel Emelyanov It's only needed to start hints via proxy, but proxy can do it without gossiper argument Closes scylladb/scylladb#19894 * github.com:scylladb/scylladb: storage_service: Remote gossiper argument from join_cluster() proxy: Use remote gossiper to start hints resource manager hints: Const-ify gossiper references and anchor pointers	2024-08-01 10:18:14 +02:00
Michael Litvak	c944e28e43	db: fix waiting for counter update operations on table stop When a table is dropped it should wait for all pending operations in the table before the table is destroyed, because the operations may use the table's resources. With counter update operations, currently this is not the case. The table may be destroyed while there is a counter update operation in progress, causing an assert to be triggered due to a resource being destroyed while it's in use. The reason the operation is not waited for is a mistake in the lifetime management of the object representing the write in progress. The commit fixes it so the object lives for the duration of the entire counter update operation, by moving it to the `do_with` list. Fixes scylladb/scylla-enterprise#4475 Closes scylladb/scylladb#19948	2024-08-01 09:39:49 +02:00
Nadav Har'El	5411559a94	test/cql-pytest: test ALLOW FILTERING in intersection of two indexes A user complained that ScyllaDB is incompatible with Cassandra when it requires ALLOW FILTERING on a restriction like WHERE x=1 AND y=1 where x and y are two columns with secondary indexes. In the tests added in this patch we show that: 1. Scylla is compatible with Cassandra when the traditional "CREATE INDEX" is used - ALLOW FILTERING is required in this case in both Cassandra and Scylla. 2. If SAI is used in Cassandra (CREATE CUSTOM INDEX USING 'SAI'), indeed ALLOW FILTERING becomes optional. I believe this is incorrect so I opened CASSANDRA-19795. These two tests combined show that we're not incompatible with Cassandra, rather Cassandra's two index implementations are incompatible between themselves, and Scylla is in fact compatible in this case with Cassadra's traditional index and not with SAI. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#19909	2024-07-31 14:01:29 +03:00
Laszlo Ersek	e67eb0ccc1	test/sstable: coroutinize do_write_sst() Make do_write_sst() easier to read by coroutinizing it. Closes #19803. Suggested-by: Benny Halevy <bhalevy@scylladb.com> Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#19937	2024-07-31 13:59:26 +03:00
Kefu Chai	020333fcf1	sstables: fix a typo in comment s/guranteed/guaranteed/ Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19946	2024-07-31 13:58:09 +03:00
Tomasz Grabiec	28de5231f4	test: perf: tablet_load_balancing: Test with higher shard and tablet counts We have up to 200 shards in production, so test this to catch performance issues.	2024-07-31 12:57:15 +02:00
Tomasz Grabiec	19b7fb3a4d	tablets: load_balancer: Avoid quadratic complexity when finding best candidate If the source and destination shards picked for migration based on global tablet balance do not have a good candidate in terms of effect on per-table balance, the algorithm explores other source shards and destinations. This has quadratic complexity in terms of shard count in the worst case, when there are no good candidates. Since we can have up to ~200 shards, this can slow down scheduling significantly. I saw total scheduling time of 5 min in the following run: scylla perf-load-balancing -c1 -m1G --iterations=8 \ --nodes=4 --tablets1=1024 --tablets2=8096 \ --rf1=2 --rf2=3 --shards=256 To improve, change the apprach to first find the best source shard and then best target shard, sequentially. So it's now linear in terms of shard count. After the change, the total scheduling time in that run is down to 4s. Minimizing source and destination metrics piece-wise minimizes the combined metric, so badness of the best candidate doesn't suffer after this change.	2024-07-31 12:57:15 +02:00
Tomasz Grabiec	93df82032f	tablets: load_balancer: Maintain load sketch properly during intra-node migration Affects only intra-node migration. The code was recording destination shard as taken and did not un-take it in case we skipped the migration due to lack of candidates. Noticed during code review. Impact is minor, since even if this leads to suboptimal balance, the next scheduling round should fix it. Also, the source shard was not unloaded, but that should have no impact on decisions. But to be future-proof, better to maintain the load accurately in case the algorithm is extended with more steps.	2024-07-31 12:57:15 +02:00
Tomasz Grabiec	88988ce0db	tablets: load_balancer: Use "drained" flag Cleanup / optimization.	2024-07-31 12:57:15 +02:00
Tomasz Grabiec	56801b7cb7	test: perf: tablet_load_balancing: Report load balancer stats	2024-07-31 12:57:15 +02:00
Tomasz Grabiec	90c9934099	tablets: load_balancer: Move load_balancer_stats_manager to header file So that stats can be accessed outside tablet allocator.	2024-07-31 12:57:15 +02:00
Anna Stuchlik	ae28880fc8	doc: enable publishing docs for branch-6.1 This commit enables publishing documentation from branch-6.1. The docs will be published as UNSTABLE (the warning about version 6.1 being unstable will be displayed). Fixes https://github.com/scylladb/scylladb/issues/19926 No backport is required. Closes scylladb/scylladb#19931	2024-07-31 12:48:51 +02:00
Kamil Braun	c05e077a13	Merge 'raft: fix the shutdown phase being stuck' from Emil Maskovsky Some of the calls inside the `raft_group0_client::start_operation()` method were missing the abort source parameter. This caused the repair test to be stuck in the shutdown phase - the abort source has been triggered, but the operations were not checking it. This was in particular the case of operations that try to take the ownership of the raft group semaphore (`get_units(semaphore)`) - these waits should be cancelled when the abort source is triggered. This should fix the following tests that were failing in some percentage of dtest runs (about 1-3 of 100): * TestRepairAdditional::test_repair_kill_1 * TestRepairAdditional::test_repair_kill_3 Fixes scylladb/scylladb#19223 Closes scylladb/scylladb#19860 * github.com:scylladb/scylladb: raft: fix the shutdown phase being stuck raft: use the abort source reference in raft group0 client interface	2024-07-31 12:10:30 +02:00
Pavel Emelyanov	93ed978729	view_builder: Drop unused members There's a counter and a shared future on board, that used to facilitate start-time barrier synchronization. Now they are not needed. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-31 12:59:40 +03:00
Pavel Emelyanov	613161c7b9	view_builder: Use cross-shard barrier on start When starting, view builder spawns an async background fibers, and upon its completion each shard needs to wait for other shards to do the same. This is exactly what cross-shard barrier is about, so instead of synchronizing via v.b.'s shard-0 instance, use the barrier. This makes the view_builder::start() shorder and earier to read. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-31 12:56:25 +03:00
Pavel Emelyanov	fb1b749445	view_builder: Add cross-shard barrier to its .start() method The barrier will be used by next patch to synchronize shards with each other. When passed to invoke_on_all() lambda like this, each lambda gets its its copy of the barrier "handler" that maintains shared state across shards. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-31 12:54:28 +03:00
Tomasz Grabiec	94cce4b7d3	tablets: load_balancer: Split evaluate_candidate() into src and dst part Those parts will be used separately later.	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	4df2abe47a	tablets: load_balancer: Optimize evaluate_candidate() Moves load computation out of the hot path by relying on data structures maintained globally during plan making.	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	5e7facd543	tablets: load_balancer: Add more statistics	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	be055977c9	tablets: load_balancer: Track load per table on cluster level	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	81fcee2040	tablets: load_balancer: Track load per table on node level	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	e7ef7419dc	tablets: load_balancer: Use a single load sketch for tracking all nodes This is code simplification and optimization. Avoids multiple passes of tablet metadata to consturct load sketch for each target node.	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	352b8e0ddd	locator: load_sketch: Introduce populate_dc()	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	9a7afd334b	tablets: load_balancer: Modify target load sketch only when emitting migration This avoids the need to unpick() a replica when the candidate is not selected. Optimization.	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	b78657ce7d	locator: load_sketch: Introduce get_most_loaded_shard()	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	de404471b7	locator: load_sketch: Introduce get_least_loaded_shard()	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	8fbfd595bb	locator: load_sketch: Optimize pick()/unload() They are executed frequently during tablet scheduling. Currently, they have time complexity of O(N*log(N)) in terms of shard count. With large shard counts, that has significant overhead. This patch optimizes them down to O(log(N)).	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	d0b0f95849	locator: load_sketch: Introduce load_type	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	8f3b623144	test: perf: tablet_load_balancing: Report total tablet counts	2024-07-31 11:38:17 +02:00
Tomasz Grabiec	662a0ff038	test: perf: tablet_load_balancing: Print run parameters in the single simulation case too	2024-07-31 11:38:16 +02:00
Tomasz Grabiec	a040404875	test: perf: tablet_load_balancing: Report time it took to schedule migrations	2024-07-31 11:38:16 +02:00
Tomasz Grabiec	ae7fd80554	tablets: load_balancer: Log table load stats after each migration	2024-07-31 11:38:16 +02:00
Tomasz Grabiec	b8996a0f59	tablets: load_balancer: Log per-shard load distribution in debug level	2024-07-31 11:38:16 +02:00
Tomasz Grabiec	469e2f3f90	tablets: load_balancer: Improve per-table balance Tablet load balancer tries to equalize tablet load between shards by moving tablets. Currently, the tablet load balancer assumes that each tablet has the same hotness. This may not be true, and some tables may be hotter than others. If some nodes end up getting more tablets of the hot table, we can end up with request load imbalance and reduced performance. In `79d0711c7e` we implemented a mitigation for the problem by randomly choosing the table whose tablet replica should be moved. This should improve fairness of movement. However, this proved to not be enough to get a good distribution of tablets. This change improves candidate selection to not relay on randomness but rather evaluating candidates with respect to the impact on load imbalance. Also, if there is no good candidate, we consider picking other source shards, not the most-loaded one. This is helpful because when finishing node drain we get just a few candidates per shard, all of which may belong to a single table, and the destination may already be overloaded with that table. Another shard may contain tablets of another table which is not yet overloaded on the destination. And shards may be of similar load, so it doesn't matter much which shard we choose to unload. We also consider other destinations, not the least-loaded one. This helps when draining nodes and the source node has few shard candidates. Shards on the destination may have similar load so there is more than one good destinatin candidate. By limiting ourselves to a single shard, we increase the chance that we're overload the table on that shard. The algorithm was evaluated using "scylla perf-load-balancing", which simulates a sequeunce of 8 node bootstraps and decommissions for different node and shard counts, RF, and tablet counts. For example, for the following parameters: params: {iterations=8, nodes=5, tablets1=128 (2.4/sh), tablets2=512 (9.6/sh), rf1=3, rf2=3, shards=32} The results are: After: Overcommit : init : {table1={shard=1.25 (best=1.25), node=1.00}, table2={shard=1.04 (best=1.04), node=1.00}} Overcommit : worst: {table1={shard=1.50 (best=1.25), node=1.02}, table2={shard=1.12 (best=1.04), node=1.01}} Overcommit : last : {table1={shard=1.25 (best=1.25), node=1.00}, table2={shard=1.04 (best=1.04), node=1.00}} Before: Overcommit (old) : init : {table1={shard=1.25 (best=1.25), node=1.00}, table2={shard=1.04 (best=1.04), node=1.00}} Overcommit (old) : worst: {table1={shard=4.00 (best=1.25), node=1.81}, table2={shard=1.25 (best=1.04), node=1.11}} Overcommit (old) : last : {table1={shard=2.50 (best=1.25), node=1.41}, table2={shard=1.25 (best=1.04), node=1.05}} So shard overcommit for table1 was reduced from 4 to 1.5. Overcommit of 4 means that the most-loaded shard has 4 times more tablets than the average per-shard load in the cluster. Also, node overcommit for table1 was reduced from 1.81 to 1.02. The magnitude of improvement depends greatly on test configurtion, so on topology and tablet distribution. The algorithm is not perfect, it finds a local optimum. In the above test, overcommit of 1.5 is not the best possible (1.25). One of the reason why the current algorithm doesn't achieve best distribution is that it works with a single movement at a time and replication constraints limit the choice of destinations. Viable destinations for remaining candidates may by only on nodes which are not least-loaded, and we won't be able to fill the least loaded node. Doing so would require more complex movement involving moving a tablet from one of the destination nodes which doesn't have a replica on the least loaded node and then replacing it with the candidate from the source node. Another limitation is that the algorithm can only fix balance by moving tablets away from most loaded nodes, and it does so due to imbalance between nodes. So it cannot fix the imbalance which is already present on the nodes if there is not much to move due to similar load between nodes. It is designed to not make the imbalance worse, so it works good if we started in a good shape. Fixes #16824	2024-07-31 11:38:16 +02:00
Tomasz Grabiec	b7661aa6c9	tablets: load_balancer: Extract check_convergence() Will be reused when evaluating different targets for migration in later stages. The refactoring drops updating of _stats.for_dc(dc).stop_no_candidates and we update _stats.for_dc(dc).stop_load_inversion in both cases where convergence check may fail. The reason is that stat updates must be outside check_convergence(), since the new use case should not update those stats (it doesn't stop balancing, just drops candidates). Propagating the information for distinguishing the two cases would be a burden. But it's not necessary, since both cases are actually load inversion cases, one pre-migration the other post-migration, so we don't need the distinction. It's actually wrong to increment stop_no_candidates, since there may still be candidates, it's the load which is inverted.	2024-07-31 11:26:11 +02:00
Tomasz Grabiec	41e643ddb9	tablets: load_balancer: Extract nodes_by_load_cmp Will be reused in a different place.	2024-07-31 11:26:11 +02:00
Tomasz Grabiec	8a7257971d	tablets: load_balancer: Maintain tablet count per table	2024-07-31 11:26:11 +02:00
Tomasz Grabiec	4e4f13ac9d	tablets: load_balancer: Reuse src_node_info	2024-07-31 11:26:11 +02:00
Tomasz Grabiec	71b8d6b7aa	test: perf: tablet_load_balancing: Print warnings about bad overcommit	2024-07-31 11:26:11 +02:00
Tomasz Grabiec	0d50a028a5	test: perf: tablet_load_balancing: Allow running a single simulation	2024-07-31 11:26:11 +02:00
Tomasz Grabiec	3f3660c3fe	test: perf: tablet_load_balancing: Report best possible shard overcommit	2024-07-31 11:26:11 +02:00
Tomasz Grabiec	c89a320925	test: perf: tablet_load_balancing: Report global shard overcommit Rather than maximum per-node shard overcommit. Global shard overcommit is a better metric since we want to equalize global load not just per-node load.	2024-07-31 11:26:11 +02:00
Emil Maskovsky	5dfc50d354	raft: fix the shutdown phase being stuck Some of the calls inside the `raft_group0_client::start_operation()` method were missing the abort source parameter. This caused the repair test to be stuck in the shutdown phase - the abort source has been triggered, but the operations were not checking it. This was in particular the case of operations that try to take the ownership of the raft group semaphore (`get_units(semaphore)`) - these waits should be cancelled when the abort source is triggered. This should fix the following tests that were failing in some percentage of dtest runs (about 1-3 of 100): * TestRepairAdditional::test_repair_kill_1 * TestRepairAdditional::test_repair_kill_3 Fixes scylladb/scylladb#19223	2024-07-31 09:18:54 +02:00
Emil Maskovsky	2dbe9ef2f2	raft: use the abort source reference in raft group0 client interface Most callers of the raft group0 client interface are passing a real source instance, so we can use the abort source reference in the client interface. This change makes the code simpler and more consistent.	2024-07-31 09:18:54 +02:00
Benny Halevy	82333036f3	cell_locker: maybe_rehash: reindent Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-31 10:06:07 +03:00
Benny Halevy	8853adea96	cell_locker: maybe_rehash: ignore allocation failures `maybe_rehash` is complimentary and is not strictly required to succeed. If it fails, it will retry on the next call, but there's no reason to throw a bad_alloc exception that will fail its caller, since `maybe_rehash` is called as the final step after the caller has already succeeded with its action. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-31 10:06:06 +03:00
Pavel Emelyanov	9214aecbe7	storage_service: Remove orphan forward declaration of a method The start_sys_dist_ks() itself was removed by `bc051387c5` Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#19928	2024-07-30 16:17:49 +03:00
Benny Halevy	e58ca8c44b	service_level_controller: stop: always call subscription on_abort We want to call `service_level_controller::do_abort()` in all cases. The current code (introduced in `535e5f4ae7`) calls do_abort if abort was not requested, however, since it does so by checking the subscription bool operator, it would miss the case where abort was already requested before the subscription took place (in service_level_controller ctor). With scylladb/seastar@470b539b1c and scylladb/seastar@8ecce18c51 we can just unconditionally call the subscription `on_abort` method, that ensures only-once semantics, even if abort was already requested at subscription time. Fixes scylladb/scylladb#19075 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#19929	2024-07-30 13:23:17 +03:00
Kefu Chai	35394c3f9a	docs/dev: fix a typo remove the extraneous "is". Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19902	2024-07-30 10:46:25 +03:00
Pavel Emelyanov	97154b0671	Merge 'mapreduce_service: complete coroutinization' from Avi Kivity mapreduce_server was previously coroutinized, but only partially. This series completes coroutinization and eliminates remaining continuation chains. None of this code is performance sensitive as it runs at the super-coordinator level and is amortized over a full scan of the entire table. No backport needed as this is a cleanup. Closes scylladb/scylladb#19913 * github.com:scylladb/scylladb: mapreduce_service: reindent mapreduce_service: coroutinize retrying_dispatcher::dispatch_to_node() mapreduce_service: coroutinize dispatch() inner lambda	2024-07-30 10:44:34 +03:00
Nadav Har'El	d293a5787f	alternator: exclude CDC log table from ListTables The Alternator command ListTables is supposed to list actual tables created with CreateTable, and should list things like materialized views (created for GSI or LSI) or CDC log tables. We already properly excluded materialized views from the list - and had the tests to prove it - but forgot both the exclusion and the testing for CDC log tables - so creating a table xyz with streams enable would cause ListTables to also list "xyz_scylla_cdc_log". This patch fixes both oversights: It adds the code to exclude CDC logs from the output of ListTables, add adds a test which reproduces the bug before this fix, and verifies the fix works. Fixes #19911. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#19914	2024-07-30 10:43:29 +03:00
Nadav Har'El	ca8b91f641	test: increase timeouts for /localnodes test In commit `bac7c33313` we introduced a new test for the Alternator "/localnodes" request, checking that a node that is still joining does not get returned. The tests used what I thought were "very high" timeouts - we had a timeout of 10 seconds for starting a single node, and injected a 20 second sleep to leave us 10 seconds after the first sleep. But the test failed in one extremely slow run (a debug build on aarch64), where starting just a single node took more than 15 seconds! So in this patch I increase the timeouts significantly: We increase the wait for the node to 60 seconds, and the sleeping injection to 120 seconds. These should definitely be enough for anyone (famous last words...). The test doesn't actually wait for these timeouts, so the ridiculously high timeouts shouldn't affect the normal runtime of this test. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#19916	2024-07-30 10:41:48 +03:00
Avi Kivity	52ee6127dd	Merge 'Use boto3 in object_store test to list bucket' from Pavel Emelyanov There's a test in object_store suite that verifies the contents of a bucket. It does with the plain http request, but unfortunately this doesn't work -- even local minio uses restricted bucket and using plain http request results in 403(Forbidden) error code. Test doesn't check it and continues working with empty list of objects which, in turn, is what it expects to see. The fix is in using boto3. With it, the acc/secret pair is picked up and listing the bucket finally works. Closes scylladb/scylladb#19889 * github.com:scylladb/scylladb: test/object_store: Use boto3.resource to list bucket test/object_store: Add get_s3_resource() helper	2024-07-29 13:49:50 +03:00
Pavel Emelyanov	8b1a106b62	test/object_store: Use boto3.resource to list bucket Instead of plain http request, use the power of boto3 package. The recently added get_s3_resource() facilitates creating one Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-29 12:29:16 +03:00
Pavel Emelyanov	172e1cb0da	test/object_store: Add get_s3_resource() helper It creates boto3.resource object that points to endpoint maintained by s3_server argument (that tests obtain via fixture). This allows using boto3 to access S3 bucket from local minio server. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-29 12:25:57 +03:00
Kefu Chai	1094c71282	cql3/statement: use compile-time format string instead of using fmt::runtime, use compile-time format string in order to detect the bad format string, or missing format arguments, or arguments which are not formattable at compile time. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19901	2024-07-28 21:54:43 +03:00
Benny Halevy	be880ab22c	Update seastar submodule * seastar 67065040...a7d81328 (30): > reactor: Initialize _aio_pollfd later > abortable_fifo: fix a typo in comment > net: Expose DNS error category > pollable_fd_state: use default-generated dtor > perftune: tune tcp_mem > scripts/perftune.py: clock source tweaking: special case Amazon and Google KVM virtualizations > abort_source: subscription: keep callback function alive after abort > github: disable ccache when building with C++ modules > github: add enable-ccache input to test.yaml > pollable_fd_state: Mark destructor protected and make non-virtual > reactor: Mark .configure() private > reactor: Set aio_nowait_supported once > reactor: Add .no_poll_aio to reactor_config > reactor: Move .max_poll_time on reactor_config > reactor: Move .task_quota on reactor_config > reactor: Move .strict_o_direct on reactor_config > reactor: Move .bypass_fsync on reactor_config > reactor: Move .max_task_backlog on reactor_config > reactor: Move .force_io_getevents_syscall on reactor_config > reactor: Move .have_aio_fsync on reactor_config > reactor: Move .kernel_page_cache on reactor_config > reactor: Move .handle_sigint on reactor_config > reactor_backend: Construct _polling_io from reactor config > reactor: Move config when constructing > reactor: Use designated initializers to set up reactor_config > native-stack: use queue::pop_eventually() in listener::accept() > abort_source: subscription: allow calling on_abort explicitly > file: document that close() returns the file object to uninitialized state > code-cleanup: do not include 'smp.hh' in 'reactor.hh' > code-cleanup: remove redundant includes of smp.hh Closes scylladb/scylladb#19912	2024-07-28 21:04:45 +03:00
Kefu Chai	36f5032b2d	db: correct the doxygen comment the parameter names do not match with the ones we are using. these comments were inherited from Origin, but we failed to update them accordingly. in this change, the comments are updated to reflect the function signatures. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19900	2024-07-28 18:24:57 +03:00
Kefu Chai	67e07bee25	build: cmake: use per-mode build dir The build_unified.sh script accepts a --build-dir option, which specifies the directory used for storing temporary files extracted from tarballs defined by the --pkgs option. When performing parallel builds of multiple modes, it's crucial that each build uses a unique build directory. Reusing the same build directory for different modes can lead to conflicts, resulting in build failures or, more seriously, the creation of tarballs containing corrupted files. so, in this change, we specify a different directory for each mode, so that they don't share the same one. Refs scylladb/scylladb#2717 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19905	2024-07-28 18:11:37 +03:00
Avi Kivity	149a47088e	mapreduce_service: reindent	2024-07-28 17:55:51 +03:00
Avi Kivity	0dd03789f3	mapreduce_service: coroutinize retrying_dispatcher::dispatch_to_node() Simplify the function by converting it to a coroutine. Note that while the final co_return co_await looks like a loop (and therefore an await would introduce an O(n) allocation), it really isn't - we retry at most once.	2024-07-28 17:54:01 +03:00
Avi Kivity	b019927a0e	mapreduce_service: coroutinize dispatch() inner lambda dispatch() is a coroutine, but the inner lambda that is executed per node is still a continuation chain. Make it uniform by converting to a coroutine.	2024-07-28 17:36:08 +03:00
Kefu Chai	ee80742c39	cql3: do not include unused headers these unused includes were identified by clangd. see https://clangd.llvm.org/guides/include-cleaner#unused-include-warning for more details on the "Unused include" warning. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19906	2024-07-28 17:29:07 +03:00
Benny Halevy	26abad23d9	sstable_directory: delete_atomically: allow sstables from multiple prefixes Currently, delete_atomically can be called with a list of sstables from mixed prefixes in two cases: 1. truncate: where we delete all the sstables in the table directory 2. tablet cleanup: similar to truncate but restricted to sstables in a single tablet replica In both cases, it is possible that sstables in staging (or quarantine) are mixed with sstables in the base directory. Until a more comprehensive fix is in place, (see https://github.com/scylladb/scylladb/pull/19555) this change just lifts the ban on atomic deletion of sstables from different prefixes, and acknowledging that the implementation is not atomic across prefixes. This is better than crashing for now, and can be backported more easily to branches that support tablets so tablet migration can be done safely in the presence of repair of tables with views. Refs scylladb/scylladb#18862 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#19816	2024-07-28 17:26:31 +03:00
Pavel Emelyanov	aaad2bbeaf	storage_service: Remote gossiper argument from join_cluster() This pointer was only needed to pull all the way down the hints resource manager start() method. It's no longer needed for that. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-26 16:29:58 +03:00
Pavel Emelyanov	a1dbaba9e1	proxy: Use remote gossiper to start hints resource manager By the time hinst resource manager is started, proxy already has its remote part initialized. Remote returns const gossiper pointer, but after previous change hints code can live with it. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-26 16:29:03 +03:00
Pavel Emelyanov	dd7c7c301d	hints: Const-ify gossiper references and anchor pointers There are two places in hints code that need gossiper: hist_sender calling gossiper::is_alive() and endpoint_downtime_not_bigger_than() helper in manager. Both can live with const gossiper, so the dependency references and anchor pointers can be restricted to const too. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-26 16:28:54 +03:00
Lakshmi Narayanan Sreethar	27b305b9d1	boost/bloom_filter_test: wait for total memory reclaimed update The testcase `test_bloom_filter_reclaim_during_reload` checks the SSTable manager's `_total_memory_reclaimed` against an expected value to verify that a Bloom filter was reloaded. However, it does not wait for the manager to update the variable, causing the check to fail if the update has not occurred yet. Fix it by making the testcase wait until the variable is updated to the expected value. Fixes #19879 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com> Closes scylladb/scylladb#19883	2024-07-26 08:15:11 +03:00
Tomasz Grabiec	851da230c8	Merge 'db/view: drop view updates to replaced node marked as left' from Piotr Dulikowski When a node that is permanently down is replaced, it is marked as "left" but it still can be a replica of some tablets. We also don't keep IPs of nodes that have left and the `node` structure for such node returns an empty IP (all zeros) as the address. This interacts badly with the view update logic. The base replica paired with the left node might decide to generate a view update. Because storage proxy still uses IPs and not host IDs, it needs to obtain the view replica's IP and tell the storage proxy to write a view update to that node - so, it chooses 0.0.0.0. Apparently, storage proxy decides to write a hint towards this address - hinted handoff on the other hand operates on host IDs and not IPs, so it attempts to translate the IP back, which triggers an assertion as there is no replica with IP 0.0.0.0. As a quick workaround for this issue just drop view updates towards nodes which seem to have IPs that are all zeros. It would be more proper to keep the view updates as hints and replay them later to the new paired replica, but achieving this right now would require much more significant changes. For now, fixing a crash is more important than keeping views consistent with base replicas. In addition to the fix, this PR also includes a regression test heavily based on the test that @kbr-scylla prepared during his investigation of the issue. Fixes: scylladb/scylladb#19439 This issue can cause multiple nodes to crash at once and the fix is quite small, so I think this justifies backporting it to all affected versions. 6.0 and 6.1 are affected. No need to backport to 5.4 as this issue only happens with tablets, and tablets are experimental there. Closes scylladb/scylladb#19765 * github.com:scylladb/scylladb: test: regression test for MV crash with tablets during decommission db/view: drop view updates to replaced node marked as left	2024-07-25 11:47:14 +02:00
Michael Litvak	6f25f4b387	mv: skip reading rows when generating partition tombstone update when deleting a base partition, in some cases we can update the view by generating a single partition deletion update, instead of generating a row deletion update for each of the partition rows. If this is the case for all the affected views, and there are no other updates besides deleting the partition, then we can skip reading and iterating over all the rows, since this won't generate any additional updates that are not covered already.	2024-07-25 11:12:58 +03:00
Michael Litvak	d0b02dc0d0	mv: delete a partition in a single operation when applicable Currently when a partition is deleted from the base table, we generate a row tombstone update for each one of the view rows in the partition. When the partition key in the view is the same as the base, maybe in a different order, this can be done more efficiently - The whole corresponding view partition can be deleted with one partition tombstone update. With this commit, when generating view updates, if the update mutation has a partition tombstone then for the views which have the same partition key we will generate a partition tombstone update, and skip the individual row tombstone updates. Fixes scylladb/scylladb#8199	2024-07-25 11:12:58 +03:00
Michael Litvak	98cc707c76	cql-pytest: move ScyllaMetrics to util file to allow reuse ScyllaMetrics is a useful generic component for retrieving metrics in a pytest. The commit moves the implementation from test_shedding.py to util.py to make it reusable in other tests in cql-pytest.	2024-07-25 11:12:58 +03:00
Botond Dénes	1bfe73c2ea	Merge 'Order API endpoints registration in main' from Pavel Emelyanov There are few api::set_foo()-s left in main that are placed in ~~random~~ legacy order. This PR fixes it and makes few more associated cleanups. refs: #2737 Closes scylladb/scylladb#19682 * github.com:scylladb/scylladb: api: Unset cache_service endpoints on stop main: Don't ignore set_cache_service() future api: Move storage API few steps above api: Register token-metadata API next to token-metadata itsels api: Do not return zero local host-id api: Move snitch API registration next to snitch itself	2024-07-25 09:59:38 +03:00
Pavel Emelyanov	456dbc122b	api: Unset cache_service endpoints on stop They currently stay registered long after the dependent services get stopped. There's a need for batch unsetting (scylladb/seastar#1620), so currently only this explicit listing :( Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 18:51:32 +03:00
Pavel Emelyanov	61fb0ad996	main: Don't ignore set_cache_service() future The call itself seem to be in wrong place -- there's no "cache service" also the API uses database and snapshot_ctl to work on. So it deserves more cleanup, but at least don't throw the returned future<> away. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 18:51:32 +03:00
Pavel Emelyanov	e1eb48f9c2	api: Move storage API few steps above The sequence currently is sharded<storage_service>.start() sharded<query_processor>.invoke_on_all(start_remote) api::set_server_storage_service() The last two steps can be safely swapped to keep storage service API next to its service. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 18:51:32 +03:00
Pavel Emelyanov	6ae09cc6bf	api: Register token-metadata API next to token-metadata itsels Right now API registration happens quite late because it waits storage service to register its "function" first. This can be done beforeheand and the t.m. API can be moved to where it should be. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 18:51:32 +03:00
Pavel Emelyanov	10566256fd	api: Do not return zero local host-id The local host id is read from local token metadata and returned to the caller as string. The t.m. itself starts with default-constructed host id vlaue which is updated later. However, even such "unset" host id value can be rendered as string without errors. This makes the correct work of the API endpoint depend on the initialization sequence which may (spoilter: it will) change in the future. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 18:51:32 +03:00
Pavel Emelyanov	29738f0cb6	api: Move snitch API registration next to snitch itself Once sharded<snitch> is started, it can register its handlers Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 18:51:07 +03:00
Pavel Emelyanov	6357755624	replica: Remove keyspace::config::datadir It's finally no longer used. Now only sstables storage code "knows" that keyspace may have its on-disk directory. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 17:45:51 +03:00
Pavel Emelyanov	f767e25c8b	sstables/storage: Evaluate path for keyspace directory in storage Currently the init_keyspace_storage() expects that the caller would tell it where the ks directory is, but it's not nice as keyspace may not necessarity keep its sstables in any directory. This patch moves the directory path evaluation into storage code, specifically to the lambda that is called for on-disk sstables. The way directory is evaluated mirrors the one from make_keyspace_config() that will be removed by next patch. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 17:45:50 +03:00
Pavel Emelyanov	3ae41bd6f6	sstables/storage: Add sstables_manager arg to init_keyspace_storage() Will be needed by next patch Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-07-24 17:41:45 +03:00
Botond Dénes	6337372b9d	test/boost/reader_concurrency_semaphore_test: un-flake test admission The admission test has a section which tests admission when the semaphore has inactive reads. This section (and therefore the enire test) became flaky lately, after a seemingly unrelated seastar upgrade, which improved timers. The cause of the flakyness is the permit which is made inactive later: this permit is created with 0 timeout (times out immediately). For some time now, when the timeout timer of a permit fires, if the permit is inactive, it is evicted. This is what makes the test fail: the inactive read times out and ends up evicting this permit, which is not expected for the test. The reason this was not a problem before, is that the test finishes very quickly, usually, before the timer could even be polled by the reactor. The recent seastar changes changed this and now the timer sometimes get polled and fires, failing the test. Fixes: #19801 Closes scylladb/scylladb#19859	2024-07-24 13:04:50 +03:00
Takuya ASADA	02b20089cb	scylla_raid_setup: install update-initramfs when it's not available scylla_raid_setup may fail on Ubuntu minimal image since it calls update-initramfs without installing. Closes scylladb/scylladb#19651	2024-07-24 11:55:16 +03:00
Pavel Emelyanov	b02d20d12d	Merge 'Minor improvements around compaction groups' from Raphael "Raph" Carvalho Minor changes, no backporting needed. Closes scylladb/scylladb#19723 * github.com:scylladb/scylladb: replica: rename for_each_const_compaction_group() replica: Fix comment about compaction group replica: remove unused compaction_group_vector	2024-07-24 11:22:24 +03:00
Nadav Har'El	edc5bca6b1	alternator: do not allow authentication with a non-"login" role Alternator allows authentication into the existing CQL roles, but roles which have the flag "login=false" should be refused in authentication, and this patch adds the missing check. The patch also adds a regression test for this feature in the test/alternator test framework, in a new test file test/alternator/cql_rbac.py. This test file will later include more tests of how the CQL RBAC commands (CREATE ROLE, GRANT, REVOKE) affect authentication and authorization in Alternator. In particular, these tests need to use not just the DynamoDB API but also CQL, so this new test file includes the "cql" fixture that allows us to run CQL commands, to create roles, to retrieve their secret keys, and so on. Fixes scylladb/scylladb#19735 Closes scylladb/scylladb#19740	2024-07-24 08:20:23 +02:00
Botond Dénes	84db147c58	Merge 'tasks: introduce virtual tasks' from Aleksandra Martyniuk Introduce virtual tasks - task manager tasks which cover cluster-wide operations. Virtual tasks aren't kept in memory, instead their statuses are retrieved from associated service when user requests them with task manager API. From API users' perspective, virtual tasks behave similarly to regular tasks, but they can be queried from any node in a cluster. Virtual tasks cannot have a parent task. They can have children on each node in a cluster, but do not keep references to them. So, if a direct child of a virtual task is unregistered from task manager, it will no longer be shown in parent's children vector. virtual_task class corresponds to all virtual tasks in one group. If users want to list all tasks in a module, a virtual_task returns all recent supported operations; if they request virtual task's status - info about the one specified operation is presented. Time to live, number of tracked operations etc. depend on the implementation of individual virtual_task. All virtual_tasks are kept only on shard 0. Refs: https://github.com/scylladb/scylladb/issues/15852 New feature, no backport needed. Closes scylladb/scylladb#16374 * github.com:scylladb/scylladb: docs: describe virtual tasks db: node_ops: filter topology request entries test: add a topology suite for testing tasks node_ops: service: create streaming tasks node_ops: register node_ops_virtual_task in task manager service: node_ops: keep node ops module in storage service node_ops: implement node_ops_virtual_task methods db: service: modify methods to get topology_requests data db: service: add request type column to topology_requests node_ops: add task manager module and node_ops_virtual_task tasks: api: add virtual task support to get_task_status_recursively tasks: api: add virtual task support tasks: api: add virtual tasks support to get_tasks tasks: add task_handler to hide task and virtual_task differences from user tasks: modify invoke_on_task tasks: implement task_manager::virtual_task::impl::get_children tasks: keep virtual tasks in task manager tasks: introduce task_manager::virtual_task	2024-07-24 08:34:28 +03:00
Botond Dénes	0bb6413ea5	Merge 'github: disable scheduled workflow on forks' from Kefu Chai as these workflows are scheduled periodically, and if they fail, notifications are sent to the repo's owner. to minimize the surprises to the contributors using github, let's disable these workflows on fork repos. Closes scylladb/scylladb#19736 * github.com:scylladb/scylladb: github: do not run clang-tidy as a cron job github: disable scheduled workflow on forks	2024-07-24 07:50:39 +03:00
Avi Kivity	3c930a61c9	Merge 'test: scylla_cluster: support more test scenarios' from Patryk Jędrzejczak We modify `ScyllaCluster.server_start` so that it changes seeds of the starting node to all currently running nodes. This allows writing tests like ```python s1 = await manager.server_add(start=False) await manager.server_add() await manager.server_start(s1.server_id) ``` However, it disallows writing tests that start multiple clusters. To fix this, we add the `seeds` parameter to `server_start`. We also improve the logic in `ScyllaCluster.add_server` to allow writing tests like ```python await manager.server_add(expected_error="...") await manager.server_add() ``` This PR only adds improvements to the `test.py` framework, no need to backport it. Closes scylladb/scylladb#19847 * github.com:scylladb/scylladb: test: scylla_cluster: improve expected_error in add_server test: scylla_cluster: support more test scenarios test: scylla_cluster: correctly change seeds in server_start	2024-07-23 22:05:31 +03:00
Patryk Jędrzejczak	02ccd2e3af	test: scylla_cluster: improve expected_error in add_server We make two changes: - we lease the IP address of a node that failed to boot because of an expected error, - we don't log "Cluster ... added ..." when a node fails to boot because of an expected error.	2024-07-23 14:35:09 +02:00
Patryk Jędrzejczak	4079cd1a7b	test: scylla_cluster: support more test scenarios Here are some examples of tests that don't work with no initial nodes, but they should work: 1. ``` await manager.server_add(expected_error="...") await manager.server_add() ``` 2. ``` await manager.servers_add(2, expected_error="...") await manager.servers_add(2) ``` 3. ``` s1 = await manager.server_add(start=False) await manager.server_start(s1.server_id, expected_error="...") await manager.server_add() ``` 4. ``` [s1, s2] = await manager.servers_add(2, start=False) await manager.server_start(s1.server_id, expected_error="...") await manager.server_start(s2.server_id, expected_error="...") await manager.servers_add(2) ``` 5. ``` s1 = await manager.server_add(start=False) await manager.server_add() await manager.server_start(s1.server_id) ``` 6. ``` [s1, s2] = await manager.servers_add(2, start=False) await manager.servers_add(2) await manager.server_start(s1.server_id) await manager.server_start(s2.server_id) ``` In this patch, we make a few improvements to make tests like the ones presented above work. I tested all the examples above manually. From now on, servers receive correct seeds if the first servers added in the test didn't start or failed to boot. Also, we remove the assertion preventing the creation of a second cluster. This assertion failed the tests presented above. We could weaken it to make these tests pass, but it would require some work. Moreover, we have tests that intentionally create two clusters. Therefore, we go for the easiest solution and accept that a single `ScyllaCluster` may not correspond to a single Scylla cluster.	2024-07-23 14:35:09 +02:00
Patryk Jędrzejczak	e196c1727e	test: scylla_cluster: correctly change seeds in server_start We change seeds in `ScyllaCluster.server_start` to all currently running nodes. The previous code only pretended that it did it. After doing this change, writing tests that create multiple clusters is impossible. To allow it, we add the `seeds` parameter to `ManagerClient.server_start`. We use it to fix and simplify the only test that creates two clusters - `test_different_group0_ids`.	2024-07-23 14:35:08 +02:00
Aleksandra Martyniuk	d04159e7de	docs: describe virtual tasks	2024-07-23 13:35:02 +02:00
Aleksandra Martyniuk	c64cb98bcf	db: node_ops: filter topology request entries system_keyspace::get_topology_request_entries returns entries for requests which are running or have finished after specified time. In task manager node ops task set the time so that they are shown for task_ttl seconds after they have finished.	2024-07-23 13:35:02 +02:00
Aleksandra Martyniuk	36b77c0592	test: add a topology suite for testing tasks Add topology_tasks test suite for testing task manager's node ops tasks. Add TaskManagerClient to topology_tasks for an easy usage of task manager rest api. Write a test for bootstrap, replace, rebuild, decommission and remove top level tasks using the above.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	a903971a74	node_ops: service: create streaming tasks Create tasks which cover streaming part of topology changes. These tasks are children of respective node_ops_virtual_task.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	63e82764e1	node_ops: register node_ops_virtual_task in task manager	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	8e56913fdf	service: node_ops: keep node ops module in storage service Keep task manager node ops module in storage service. It will be used to create and manage tasks related to topology changes. The module is created and registered in storage service constructor. In storage_service::stop() the module is stopped and so all the remaining tasks would be unregistered immediately after they are finished.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	b97a348361	node_ops: implement node_ops_virtual_task methods	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	94282b5214	db: service: modify methods to get topology_requests data Modify get_topology_request_state (and wait_for_topology_request_completion), so that it doesn't call on_internal_error when request_id isn't in the topology_requests table if require_entry == false. Add other methods to get topology request entry.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	880058073b	db: service: add request type column to topology_requests topology_requests table will be used by task manager node ops tasks, but it loses info about request type, which is required by tasks. Add request_type column to topology_requests.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	91fbfbf98a	node_ops: add task manager module and node_ops_virtual_task Add task manager node ops module and node_ops_virtual_task. Some methods will be implemented in later patches.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	d2e6010670	tasks: api: add virtual task support to get_task_status_recursively Virtual tasks are supported by get_task_status_recursively. Currently only local descendants' statuses are shown.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	5f7f403a15	tasks: api: add virtual task support Virtual tasks are supported by get_task_status, abort_task and wait_task. Task status returned by get_task_status and wait_task: - contains task_kind to indicate whether it's virtual (cluster) or regular (node) task; - children list apart from task_id contains node address of the task.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	20ba7ceff9	tasks: api: add virtual tasks support to get_tasks task_manager/list_module_tasks/{module} starts supporting virtual tasks, which means that their stats will also be shown for users. Additional task_kind param is added to indicate whether the task is virutal (cluster-wide) or regular (node-wide). Support in other paths will be added in following patches.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	1d85b319e0	tasks: add task_handler to hide task and virtual_task differences from user Contrary to regular tasks, which are per-operation, virtual tasks are associated with the whole group of operations. There may be many operations of each group performed at the same time. Info about each running operation will be shown to a user through the API. For virtual tasks, task manager imitates a regular task covering each operation, but task_manager::tasks aren't actually created in the memory. Instead, information (e.g. status) about the operation is retrieved from associated service and passed to a user. To hide most of the differences from user, task_handler class is created. Task handler performs appropriate actions depending on task's kind. However, users need to stay conscious about the kind of task, because: - get_task_status and wait_task do not unregister virtual tasks; - time for which a virtual tasks stays in task manager depends on associated service and tasks' implementation; - number of virtual task's children shown by get_tasks doesn't have to be monotonous. API is modified to use task_handler. API-specific classes are moved to task_handler.{cc,hh}.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	abde7ba271	tasks: modify invoke_on_task Modify task_manager::invoke_on_task to also check virtual tasks.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	6029936665	tasks: implement task_manager::virtual_task::impl::get_children Return a vector of task_identity of all children of a virtual task in a cluster.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	9de8d4b5b0	tasks: keep virtual tasks in task manager Virtual tasks are kept in task manager together with regular tasks. All virtual tasks are stored on shard 0. task_manager::module::make_task is modified to consider virtual tasks as possible parents.	2024-07-23 13:35:01 +02:00
Aleksandra Martyniuk	00cfc49d18	tasks: introduce task_manager::virtual_task A virtual task is a new kind of task supported by task manager, which covers cluster-wide operations. From users' perspective virtual tasks behave similarly to task_manager::tasks. The API side of virtual tasks will be covered in the following patches. Contrary to task_manager::task, virtual task does not update its fields proactively. Moreover, no object is kept in memory for each individual virtual task's operation. Instead a service (or services) is queried on API user's demand to learn about the status of running operation. Hence the name. task_manager::virtual_task is responsible for a whole group of virtual tasks, i.e. for tracking and generating statuses of all operations of similar type. To enable tracking of some kind of operations, one needs to override task_manager::virtual_task::impl and provide implementations of the methods returning appropriate information about the operations. task_manager::virtual_task must be kept on shard 0. Similarly to task_manager::tasks, virtual tasks can have child tasks, responsible for tracking suboperations' progress. But virtual tasks cannot have parents - they are always roots in task trees. Some methods and structs will be implemented in later patches.	2024-07-23 13:35:01 +02:00
Nadav Har'El	bac7c33313	alternator: fix "/localnodes" to not return nodes still joining Alternator's "/localnodes" HTTP request is supposed to return the list of nodes in the local DC to which the user can send requests. The existing implementation incorrectly used gossiper::is_alive() to check for which nodes to return - but "alive" nodes include nodes which are still joining the cluster and not really usable. These nodes can remain in the JOINING state for a long time while they are copying data, and an attempt to send requests to them will fail. The fix for this bug is trivial: change the call to is_alive() to a call to is_normal(). But the hard part of this test is the testing: 1. An existing multi-node test for "/localnodes" assummed that right after a new node was created, it appears on "/localnodes". But after this patch, it may take a bit more time for the bootstrapping to complete and the new node to appear in /localnodes - so I had to add a retry loop. 2. I added a test that reproduces the bug fixed here, and verifies its fix. The test is in the multi-node topology framework. It adds an injection which delays the bootstrap, which leaves a new node in JOINING state for a long time. The test then verifies that the new node is alive (as checked by the REST API), but is not returned by "/localnodes". 3. The new injection for delaying the bootstrap is unfortunately not very pretty - I had to do it in three places because we have several code paths of how bootstrap works without repair, with repair, without Raft and with Raft - and I wanted to delay all of them. Fixes #19694. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes scylladb/scylladb#19725	2024-07-23 13:51:16 +03:00
Pavel Emelyanov	65565a56c3	Merge 's3/client: add client::upload_file()' from Kefu Chai this member function prepares for the backup feature, where the object to be stored in the object storage is already persisted as a file on local filesystem. this brings us two benefits: - with the file, we don't need to accumulate the payloads in memory and send them in batch, as we do in upload_sink and in upload_jumbo_sink. this puts less pressure on the memory subsystem. - with the file, we can read multiple parts in parallel if multpart upload applies to it, this helps to improve the throughput. so, this new helper is introduced to help upload an sstable from local filesystem to the object storage. Fixes https://github.com/scylladb/scylladb/issues/16287 Closes scylladb/scylladb#16387 * github.com:scylladb/scylladb: s3/client: add client::upload_file() s3/client: move constants related to aws constraints out	2024-07-23 12:39:27 +03:00
Kefu Chai	061def001d	s3/client: add client::upload_file() this member function prepares for the backup feature, where the object to be stored in the object storage is already persisted as a file on local filesystem. this brings us two benefits: - with the file, we don't need to accumulate the payloads in memory and send them in batch, as we do in upload_sink and in upload_jumbo_sink. this puts less pressure on the memory subsystem. - with the file, we can read multiple parts in parallel if multpart upload applies to it, this helps to improve the throughput. so, this new helper is introduced to help upload an sstable from local filesystem to the object storage. Fixes #16287 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-07-23 14:39:30 +08:00
Kefu Chai	6701ce50a5	s3/client: move constants related to aws constraints out minimum_part_size and aws_maximum_parts_in_piece are AWS S3 related constraints, they can be reused out of client::upload_sink and client::upload_jumbo_sink, so in this change * extract them out. * use the user-defined literal with IEC prefix for better readablity to define minimum_part_size * add "aws_" prefix to `minimum_part_size` to be more consistent. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-07-23 14:33:54 +08:00
Takuya ASADA	c3bea539b6	dist: support nonroot and offline mode for scylla-housekeeping Introduce support nonroot and offline mode for scylla-housekeeping. Closes #13084 Closes scylladb/scylladb#13088	2024-07-23 07:57:32 +03:00
Aleksandra Martyniuk	dfe3af40ed	test: tasks: adjust tests to new wait_task behavior After `c1b2b8cb2c` /task_manager/wait_task/ does not unregister tasks anymore. Delete the check if the task was unregistered from test_task_manager_wait. Check task status in drain_module_tasks to ensure that the task is removed from task manager. Fixes: #19351. Closes scylladb/scylladb#19834	2024-07-22 18:24:54 +03:00
Nadav Har'El	9eb47b3ef0	Merge 'config: round-trip boolean configuration variables' from Avi Kivity When you SELECT a boolean from system.config, it reads as true/false, but this isn't accepted on UPDATE (instead, we accept 1/0). This is surprising and annoying, so accept true/false in both directions. Not a regression, so a backport isn't strictly necessary. Closes scylladb/scylladb#19792 * github.com:scylladb/scylladb: config: specialize from-string conversion for bool config: wrap boost::lexical_cast<> when converting from strings	2024-07-22 17:53:02 +03:00
Botond Dénes	d3135db457	Merge 'commitlog: Add optional max lifetime parameter to cl instance' from Calle Wilund If set, any remaining segment that has data older than this threshold will request flushing, regardless of data pressure. I.e. even a system where nothing happends will after X seconds flush data to free up the commit log. Related to #15820 The functionality here is to prevent pathological/test cases where a silent system cannot fully process stuff like compaction, GC etc due to things like CL forcing smaller GC windows etc. Closes scylladb/scylladb#15971 * github.com:scylladb/scylladb: commitlog: Make max data lifetime runtime-configurable db::config: Expose commitlog_max_data_lifetime_in_s parameter commitlog: Add optional max lifetime parameter to cl instance	2024-07-22 17:21:33 +03:00
Botond Dénes	3ff33e9c70	Update ./tools/java submodule * ./tools/java dbaf7ba7...0b4accdd (1): > cassandra-stress: Make default repl. strategy NetworkTopologyStrategy Closes scylladb/scylladb#19818	2024-07-22 17:12:09 +03:00
Kamil Braun	8ec90a0e60	docs: extend "forbidden operations" section for Raft-topology upgrade The Raft-topology upgrade procedure must not be run concurrently with version upgrade. Closes scylladb/scylladb#19746	2024-07-22 12:45:38 +03:00
Botond Dénes	591876b44e	Merge 'sstables: do not reload components of unlinked sstables' from Lakshmi Narayanan Sreethar The SSTable is removed from the reclaimed memory tracking logic only when its object is deleted. However, there is a risk that the Bloom filter reloader may attempt to reload the SSTable after it has been unlinked but before the SSTable object is destroyed. Prevent this by removing the SSTable from the reclaimed list maintained by the manager as soon as it is unlinked. The original logic that updated the memory tracking in `sstables_manager::deactivate()` is left in place as (a) the variables have to be updated only when the SSTable object is actually deleted, as the memory used by the filter is not freed as long as the SSTable is alive, and (b) the `_reclaimed.erase(sst)` is still useful during shutdown, for example, when the SSTable is not unlinked but just destroyed. Fixes https://github.com/scylladb/scylladb/issues/19722 Closes scylladb/scylladb#19717 github.com:scylladb/scylladb: boost/bloom_filter_test: add testcase to verify unlinked sstables are not reloaded sstables: do not reload components of unlinked sstables sstables/sstables_manager: introduce on_unlink method	2024-07-22 12:08:25 +03:00
Avi Kivity	358147959e	Merge 'keep table directory open for flushing' from Laszlo Ersek `filesystem_storage` methods frequently call `sync_directory()`, for the sake of flushing (sync'ing) a directory. `sync_directory()` always brackets the sync with open and close, and given that most `sync_directory()` calls target the sstable base directory, those repeated opens and closes are considered wasteful. Rework the `filesystem_storage::_dir` member (from a mere pathname) so that it stand for an `opened_directory` object, which keeps the sstable base directory open, for the purpose of repeated sync'ing. Resolves #2399. Closes scylladb/scylladb#19624 * github.com:scylladb/scylladb: sstables/storage: synch "dst_dir" more leanly in create_links_common() sstables/storage: close previous directory asynchronously upon dir change sstables/storage: futurize change_dir_for_test() sstables/storage: sync through "opened_directory" in filesystem...::move() sstables/storage: sync through "opened_directory" in the "easy" cases sstables/storage: introduce "opened_directory" class	2024-07-21 17:07:44 +03:00
Yaron Kaikov	d3cbe04130	.github/mergify.yml: update conf to support `6.1` Modify Mergify configuation to support `6.1` instead of `5.2` which is EOL Closes scylladb/scylladb#19810	2024-07-21 17:02:19 +03:00
Łukasz Paszkowski	781eb7517c	api/system: add highest_supported_sstable_format path Current upgrade dtest rely on a ccm node function to get_highest_supported_sstable_version() that looks for r'Feature (.*)_SSTABLE_FORMAT is enabled' in the log files. Starting from scylla-6.0 ME_SSTABLE_FORMAT is enabled by default and there is no cluster feature for it. Thus get_highest_supported_sstable_version() returns an empty list resulting in the upgrade tests failures. This change introduces a seperate API path that returns the highest supported sstable format (one of la, mc, md, me) by a scylla node. Fixes scylladb/scylladb#19772 Backports to 6.0 and 6.1 required. The current upgrade test in dtest checks scylla upgrades up to version 5.4 only. This patch is a prerequisite to backport the upgrade tests fix in dtest. Closes scylladb/scylladb#19787	2024-07-21 17:00:19 +03:00
Avi Kivity	36b57f3432	Merge 'token: inline optimizations' from Benny Halevy This series contains several optimizations for dht::token around its comparison functions as well as minimum_token and maximum_token definitions, by moving them inline into dht/token.hh This results in a nice improvement in perf-simple-query: ``` ==> perf-simple-query.pre <== (`21c67a5a64`) throughput: mean=95774.01 standard-deviation=1129.83 median=96243.64 median-absolute-deviation=1090.08 maximum=96864.09 minimum=94471.19 instructions_per_op: mean=41813.68 standard-deviation=16.27 median=41809.29 median-absolute-deviation=7.02 maximum=41841.64 minimum=41799.41 cpu_cycles_per_op: mean=22383.19 standard-deviation=331.01 median=22254.53 median-absolute-deviation=332.26 maximum=22744.11 minimum=21996.73 ==> perf-simple-query.post.0 <== (token: move ordering operator inline) throughput: mean=96350.01 standard-deviation=640.10 median=96228.88 median-absolute-deviation=621.45 maximum=96988.16 minimum=95478.51 instructions_per_op: mean=41627.13 standard-deviation=37.55 median=41627.06 median-absolute-deviation=2.43 maximum=41679.44 minimum=41573.31 cpu_cycles_per_op: mean=22184.65 standard-deviation=151.03 median=22163.05 median-absolute-deviation=120.83 maximum=22348.49 minimum=21967.30 ==> perf-simple-query.post.1 <== (token: operator<=>: optimize the common case) throughput: mean=96778.29 standard-deviation=1719.34 median=97021.72 median-absolute-deviation=1059.56 maximum=98300.99 minimum=93893.75 instructions_per_op: mean=41590.25 standard-deviation=5.53 median=41589.50 median-absolute-deviation=4.17 maximum=41598.39 minimum=41584.57 cpu_cycles_per_op: mean=22135.33 standard-deviation=471.98 median=21969.30 median-absolute-deviation=244.89 maximum=22905.24 minimum=21685.33 ==> perf-simple-query.post.3 <== (token: always initialize data member) throughput: mean=98264.33 standard-deviation=998.49 median=98533.02 median-absolute-deviation=780.45 maximum=99075.40 minimum=96656.51 instructions_per_op: mean=41657.61 standard-deviation=22.53 median=41648.49 median-absolute-deviation=12.89 maximum=41696.81 minimum=41642.07 cpu_cycles_per_op: mean=21808.57 standard-deviation=93.63 median=21794.56 median-absolute-deviation=75.41 maximum=21949.46 minimum=21719.55 ==> perf-simple-query.post.4 <== (token: constexpr ctors, methods, and minimum/maximum_token) throughput: mean=98095.05 standard-deviation=1333.32 median=98930.22 median-absolute-deviation=906.80 maximum=99209.38 minimum=96194.25 instructions_per_op: mean=41572.28 standard-deviation=6.04 median=41574.49 median-absolute-deviation=4.76 maximum=41579.56 minimum=41564.72 cpu_cycles_per_op: mean=21831.35 standard-deviation=169.56 median=21732.86 median-absolute-deviation=102.93 maximum=22091.66 minimum=21689.63 ==> perf-simple-query.post.5 <== (token: initialize non-key tokens with min() value) throughput: mean=99502.32 standard-deviation=1003.70 median=99744.03 median-absolute-deviation=388.87 maximum=100482.95 minimum=97813.42 instructions_per_op: mean=41593.48 standard-deviation=17.27 median=41585.25 median-absolute-deviation=8.46 maximum=41619.41 minimum=41575.86 cpu_cycles_per_op: mean=21545.90 standard-deviation=86.66 median=21578.01 median-absolute-deviation=43.17 maximum=21612.41 minimum=21395.42 ``` Optimization only. No backport required Closes scylladb/scylladb#19782 * github.com:scylladb/scylladb: token: initialize non-key tokens with min() value token: make kind-based ctor private token: constexpr ctors, methods, and minimum/maximum_token token: always initialize data member everywhere: use dht::token is_{minimum,maximum} token: operator<=>: optimize the common case token: move ordering operator inline partitioner_test: add more token-level tests	2024-07-21 15:07:36 +03:00
Benny Halevy	365e1fb1b9	token: initialize non-key tokens with min() value We already have code to return min() for the minimum and maximum tokens in long_token() and raw(), so instead of using code to return it, just make sure to set it in the _data member. Note that although this change affect serialization, the existing codebase ignores the deserialized bytes and places a constant (0 before this patch, or min() with it) in _data for non-key (minumum or maximum) tokens. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:42 +03:00
Benny Halevy	9f05072527	token: make kind-based ctor private Users outside of the token module don't need to mess with the token::kind. They can only create key tokens. Never, minimum or maximum tokens, with a particular datya value. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:42 +03:00
Benny Halevy	6806112189	token: constexpr ctors, methods, and minimum/maximum_token sizeof(dht::token) is only 16 bytes and therefore it can be passed with 2 registers. There is no sense in defining minimum_token and maximum_token out of line, returning a token& to statically allocated values that require memory access/copy, while the only call sites that needs to point to the static min/max tokens are in dht::ring_position_view. Instead, they can be defined inline as constexpr functions and return their const values. Respectively, define token ctors and methods as constexpr where applicable (and noexcept while at it where applicable) Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:42 +03:00
Benny Halevy	e509ccd184	token: always initialize data member Make sure to always initalize the _data member to 0 for non-key (minimum or maximum) tokens. This allows to simplify the equality operator that now doesn't need to rely on `operator<=>` Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:42 +03:00
Benny Halevy	850f298ccd	everywhere: use dht::token is_{minimum,maximum} The is_minimum/is_maximum predicates are more efficient than comparing the the m{minimum,maximum}_token values, respectrively. since the is_* functions need to check only the token kind. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:42 +03:00
Benny Halevy	5a60ba5c5f	token: operator<=>: optimize the common case Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:42 +03:00
Benny Halevy	adc1d7f68f	token: move ordering operator inline Token comparisons are abundant. The equality operator is defined inline in dht/token.hh by calling `t1 <=> t2`, and so is `tri_compare_raw`, which `operator<=>` calls in the common path, but `operator<=>` itself is defined out of line, losing the benefits of inlining. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:42 +03:00
Benny Halevy	7e745d31ed	partitioner_test: add more token-level tests Before changing how minimum and maximum tokens are represented in memory. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-07-20 21:21:37 +03:00
Kamil Braun	ad68a7f799	Merge 'test: raft: fix the flaky `test_raft_recovery_stuck`' from Emil Maskovsky Use the rolling restart to avoid spurious driver reconnects. This can be eventually reverted once the scylladb/python-driver#295 is fixed. Fixes scylladb/scylladb#19154 Closes scylladb/scylladb#19771 * github.com:scylladb/scylladb: test: raft: fix the flaky `test_raft_recovery_stuck` test: raft: code cleanup in `test_raft_recovery_stuck`	2024-07-19 19:34:43 +02:00
Piotr Dulikowski	4571262e46	Merge 'Improve constness of functions schema code' from Marcin Maliszkiewicz In v4 of scylladb/scylladb#19598 the last commit of the patch was replaced but this change missed merge so submitting it in a separate patch. In the current patch, the original functions class correctly marks methods as const where appropriate, and the instance() method now returns a const object. This ensures protection against accidental modifications, as all changes must go through the change_batch object. Since the functions_changer class was intended to serve the same purpose, it is now redundant. Therefore, we are reverting the commit that introduced it. Relates scylladb/scylladb#19153 Closes scylladb/scylladb#19647 * github.com:scylladb/scylladb: cql3: functions: replace template with std::function in with_udf_iter() cql3: functions: improve functions class constness handling Revert "cql3: functions: make modification functions accessible only via batch class"	2024-07-19 19:23:11 +02:00
Emil Maskovsky	9ab25e5cbf	test: raft: replace the use of read_barrier work-around Replaced the old `read_barrier` helper from "test/pylib/util.py" by the new helper from "test/pylib/rest_client.py" that is calling the newly introduced direct REST API. Replaced in all relevant tests and decommissioned the old helper. Introduced a new helper `get_host_api_address` to retrieve the host API address - which in come cases can be different from the host address (e.g. if the RPC address is changed). Fixes: scylladb/scylladb#19662 Closes scylladb/scylladb#19739	2024-07-19 19:20:44 +02:00
Laszlo Ersek	680403d2cd	sstables/storage: synch "dst_dir" more leanly in create_links_common() filesystem_storage::create_links_common() runs on directories that generally differ from "_dir", thus, we can't replace its sync_directory() calls with _dir.sync(). We can still use a common (temporary) "opened_directory" object for synching "dst_dir" three times, saving two open and two close operations. This patch is best viewed with "git show -W". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:46:31 +02:00
Laszlo Ersek	0057ee2431	sstables/storage: close previous directory asynchronously upon dir change In "filesystem_storage", change_dir_for_test() and move() replace "_dir" with "opened_directory(new_dir)" using the move assignment operator. Consequently, the file descriptor underlying "_dir" is closed synchronously as a part of object destruction. Expose the async file::close() function through "opened_directory". Introduce filesystem_storage::change_dir() as a common async workhorse for both change_dir_for_test() and move(). In change_dir(), close the old directory asynchronously. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:43:19 +02:00
Laszlo Ersek	6711574646	sstables/storage: futurize change_dir_for_test() Currently change_dir_for_test() is synchronous. Make it return a future, so that we can use async operations in change_dir_for_test() overrides. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:43:19 +02:00
Laszlo Ersek	ef446c4da0	sstables/storage: sync through "opened_directory" in filesystem...::move() Near the end of filesystem_storage::move(), we sync both the old directory, and the new directory, if "delay_commit" is null. At that point, the new directory is just "_dir"; call _dir.sync() instead of sync_directory(). This patch is best viewed with "git show -W". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:14:46 +02:00
Laszlo Ersek	4d33640481	sstables/storage: sync through "opened_directory" in the "easy" cases Replace sst.sstable_write_io_check(sync_directory, _dir.native()) with _dir.sync(sst._write_error_handler) Also replace the explicit (but still relatively "easy") open_checked_directory() + flush() + flush() operations in filesystem_storage::seal() with two _dir.sync() calls. Because filesystem_storage::create_links_common() is marked "const", we need to declare "_dir" mutable. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:14:46 +02:00
Laszlo Ersek	2c01171a4d	sstables/storage: introduce "opened_directory" class "filesystem_storage::_dir" is currently of type "std::filesystem::path". Introduce a new class called "opened_directory", and change the type of "_dir" to the new class "opened_directory". "opened_directory" keeps the directory open, and offers synchronization on that open directory (i.e., without having to reopen the directory every time). In subsequent patches, that will be put to use. The opening and closing of the wrapped directory cannot easily be handled explicitly in the "filesystem_storage" member functions. ( Namely, test::store() and test::rewrite_toc_without_scylla_component() -- both in "test/lib/sstable_utils.hh" -- perform "open -> ... -> seal" sequences, and such a sequence may be executed repeatedly. For example, sstable_directory_shared_sstables_reshard_correctly() [test/boost/sstable_directory_test.cc] does just that; it "reopens" the "filesystem_storage" object repeatedly. ) Rather than trying to restrict the order of "filesystem_storage" member function calls, replace the "opened_directory" object with a new one whenever the directory pathname is re-set; namely in filesystem_storage::change_dir_for_test() and filesystem_storage::move(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:14:46 +02:00
Piotr Dulikowski	204a479e82	Merge 'db/hints: Test `manager::too_many_in_flight_hints_for()`' from Dawid Mędrek In `6e79d64`, the behavior of `manager::too_many_in_flight_hints_for()` was accidentally modified. It remained unnoticed for some time and then fixed. In this commit, we add a test verifying that the concurrency of hints being written to disk is indeed limited and the limitations are imposed properly. Refs scylladb/scylladb#17636 Fixes scylladb/scylladb#17660 Closes scylladb/scylladb#19741 * github.com:scylladb/scylladb: db/hints: Verify that Scylla limits the concurrency of written hints db/hints: Coroutinize `hint_endpoint_manager::store_hint()` db/hints: Move a constant value to the TU it's used in	2024-07-19 13:26:34 +02:00
Lakshmi Narayanan Sreethar	0615c8a46b	boost/bloom_filter_test: add testcase to verify unlinked sstables are not reloaded Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-19 13:15:57 +05:30
Lakshmi Narayanan Sreethar	31ff69a13c	sstables: do not reload components of unlinked sstables The SSTable is removed from the reclaimed memory tracking logic only when its object is deleted. However, there is a risk that the Bloom filter reloader may attempt to reload the SSTable after it has been unlinked but before the SSTable object is destroyed. Prevent this by removing the SSTable from the reclaimed list maintained by the manager as soon as it is unlinked. The original logic that updated the memory tracking in `sstables_manager::deactivate()` is left in place as (a) the variables have to be updated only when the SSTable object is actually deleted, as the memory used by the filter is not freed as long as the SSTable is alive, and (b) the `_reclaimed.erase(*sst)` is still useful during shutdown, for example, when the SSTable is not unlinked but just destroyed. Fixes #19722 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-19 13:15:57 +05:30
Lakshmi Narayanan Sreethar	dbf22848a8	sstables/sstables_manager: introduce on_unlink method Added a new method, on_unlink() to the sstable_manager. This method is now used by the sstable to notify the manager when it has been unlinked, enabling the manager to update its bookkeeping as required. The on_unlink method doesn't do anything yet but will be updated by the next patch. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-19 13:15:55 +05:30
Kefu Chai	c52f49facb	build: cmake: do not mark cqlsh noarch in `3c7af287`, cqlsh's reloc package was marked as "noarch", and its filename was updated accordingly in `configure.py`, so let's update the CMake building system accordingly. this change should address the build failure of ``` 08:48:14 [3325/4124] Generating ../Debug/dist/tar/scylla-cqlsh-6.1.0~dev-0.20240629.60955ead75ef.noarch.tar.gz 08:48:14 FAILED: Debug/dist/tar/scylla-cqlsh-6.1.0~dev-0.20240629.60955ead75ef.noarch.tar.gz /jenkins/workspace/scylla-master/scylla-ci/scylla/build/Debug/dist/tar/scylla-cqlsh-6.1.0~dev-0.20240629.60955ead75ef.noarch.tar.gz 08:48:14 cd /jenkins/workspace/scylla-master/scylla-ci/scylla/build/dist && /usr/bin/cmake -E copy /jenkins/workspace/scylla-master/scylla-ci/scylla/tools/cqlsh/build/scylla-cqlsh-6.1.0~dev-0.20240629.60955ead75ef.noarch.tar.gz /jenkins/workspace/scylla-master/scylla-ci/scylla/build/Debug/dist/tar/scylla-cqlsh-6.1.0~dev-0.20240629.60955ead75ef.noarch.tar.gz 08:48:14 Error copying file "/jenkins/workspace/scylla-master/scylla-ci/scylla/tools/cqlsh/build/scylla-cqlsh-6.1.0~dev-0.20240629.60955ead75ef.noarch.tar.gz" to "/jenkins/workspace/scylla-master/scylla-ci/scylla/build/Debug/dist/tar/scylla-cqlsh-6.1.0~dev-0.20240629.60955ead75ef.noarch.tar.gz". ``` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19710	2024-07-19 08:00:17 +03:00
Kefu Chai	34bf10050b	build: cmake: bump up the minimal required fmt to 10.0.0 in `cccec07581`, we started using a featured introduced by {fmt} v10. so we need to bump up the required version in CMake as well. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19709	2024-07-19 07:58:31 +03:00
Botond Dénes	79567c1c98	scripts/open-coredump.sh: allow complete bypass of S3 server In some cases, the S3 server will not know about a certain build and any attempt to open a coredump which was generated by this build will fail, because the S3 server returns an empty/illegal response. There is already a bypass for missing package-url in the S3 server response, but this doesn't help in the case when the response is also missing other metadata, like build-id and version info. Extend this existig mechanism with a new --scylla-package-url flag, which provides complete bypass. When provided, the S3 server will not be queried at all, instead the package is downloaded from the link and version metadata is extracted from the package itself. Closes scylladb/scylladb#19769	2024-07-18 21:43:53 +03:00
Avi Kivity	58a8fd6f19	Update tools/python3 submodule (install umask, selinux) * tools/python3 18fa79e...fbf12d0 (1): > install.sh: fix incorrect permission on strict umask Ref https://github.com/scylladb/scylladb/issues/8589 Ref https://github.com/scylladb/scylladb/issues/19775	2024-07-18 21:36:50 +03:00
Avi Kivity	7984e595ce	Update tools/java submodule (install selinux context) * tools/java 33938ec16f...dbaf7ba7db (1): > install.sh: apply correct security context on offline installer Ref https://github.com/scylladb/scylladb/issues/8589	2024-07-18 21:03:32 +03:00
Kefu Chai	4fbfecbb3e	Update seastar submodule * seastar 908ccd93...67065040 (44): > metrics: Use this_shard_id unconditionally > sstring: prevent fmt from formatting sstring as a sequence > coding style: allow lines up to 160 chars in length > src/core: remove unnecessary includes > when_all: stop using deprecated std::aligned_union_t > reactor: respect preempt requests in debug mode > core: fix -Wunused-but-set-variable > gate: add try_hold > sstring: declare nested type with typename > rpc: pass start time to `wait_for_reply()` which accepts `no_wait_type` > scripts/perftune.py: get rid of "SyntaxWarning: invalid escape sequence" > scripts/perftune.py: add support for tweaking VLAN interfaces > scripts/perftune.py: improve discovery of bond device slaves > scripts/perftune.py: refactor __learn_slaves() function > code-cleanup: add missing header guards > code-cleanup: remove redundant includes of 'reactor.hh' > code-cleanup: explicitly depend on io_desc.hh > scripts/perftune.py: aRFS should be disabled by default in non-MQ mode > code-cleanup: remove unneeded includes of fair_queue.hh > docker: fix mount of install-dependencies > code-cleanup: remove redundant includes of linux-aio.hh > fstream: reformat the doxygen comment of make_file_input_stream() > iostream: use new-style consumer to implement copy() > stall-analyser: use 0 for default value of --minimum > reactor: fix crash during metrics gathering > build: run socket test with linux-aio reactor backend > test: Add testing of connect()-ion abort ability > linux_perf_event: exclude_idle only on x86_64 > linux_perf_event: add make_linux_perf_event > stall-analyser: gracefully handle empty input > shared_token_bucket: resolve FIXME > io_tester: ensure that file object is valid when closing it > tutorial.md: fix typo in Dan Kegel's name > test,rpc: Extend simple ping-pong case > rpc: Calculate delay and export it via metrics > rpc: Exchange handler duration with server responses > rpc: Track handler execution time > rpc: Fix hard-coded constants when sending unknown verb reply > reactor: Unfriend alien and smp queues > reactor: Add and use stopped() getter > reactor: Generalize wakeup() callers > file: Use lighter access to map of fs-info-s > file: Fix indentation after previous patch > file: Don't return chain of ready futures from make_file_impl Closes scylladb/scylladb#19780	2024-07-18 20:00:15 +03:00
Avi Kivity	f7e24cf0b1	Update tools/jmx submodule (umask fix) * tools/jmx 3328a22...89308b7 (1): > install.sh: fix incorrect permission on strict umask Ref scylladb/scylladb#14383 Ref scylladb/scylladb#8589	2024-07-18 19:37:57 +03:00
Avi Kivity	c3b9e64713	Merge 'sstable::open_sstable: pass origin from the writer' from Lakshmi Narayanan Sreethar Pass origin when opening the sstable from the writer and store it in the sstable object. This will make the origin available for the entire write path. Closes scylladb/scylladb#19721 * github.com:scylladb/scylladb: sstables: use _origin in write path sstable::open_sstable: pass and store origin	2024-07-18 19:30:32 +03:00
Avi Kivity	926a02451e	Merge 'sstables/index_reader: abort reading during shutdown' from Lakshmi Narayanan Sreethar This PR adds support for aborting index reads from within `index_consume_entry_context::consume_input` when the server is being stopped. The abort source is now propagated down to the `index_consume_entry_context`, making it available for `consume_input` to check if an abort has been requested. If an abort is detected, `consume_input` will throw an exception to stop the index read operation. Closes scylladb/scylladb#19453 * github.com:scylladb/scylladb: test/boost: test abort behaviour during index read sstables/index_reader: stop consuming index when abort has been requested sstables::index_consume_entry_context: store abort_source sstable: drop old filter only after the new filter is built during rebuild sstables/sstables_manager: store abort_source in sstable_manager replica/database: pass abort_source to database constructor	2024-07-18 19:26:22 +03:00
Avi Kivity	0780228aa2	config: specialize from-string conversion for bool The yaml/json representation for bool is true/false, but boost::lexical_cast is 1/0. Specialize bool conversion to accept true/false (for yaml/json compatibilty) and 1/0 (for backward compatibility). This provides round-trip conversion for bool configs in system.config.	2024-07-18 18:38:22 +03:00
Avi Kivity	33eaa61cdd	config: wrap boost::lexical_cast<> when converting from strings Configuration uses boost::lexical_cast to convert strings to native values (e.g. bools/ints). However, boost::lexical_cast doesn't recognize true/false for bool. Since we can't change boost::lexical_cast, replace it with a wrapper that forwards directly to boost::lexical_cast. In the next step, we'll specialize it for bool.	2024-07-18 18:38:19 +03:00
Piotr Dulikowski	5ec8c06561	test: regression test for MV crash with tablets during decommission Regression test for scylladb/scylladb#19439. Co-authored-by: Kamil Braun <kbraun@scylladb.com>	2024-07-18 16:00:26 +02:00
Anna Mikhlin	cd007123c3	Update ScyllaDB version to: 6.2.0-dev	2024-07-18 16:07:07 +03:00
Avi Kivity	47e99f4e04	Merge 'Fix lwt semaphore guard accounting' from Gleb Natapov Currently the guard does not account correctly for ongoing operation if semaphore acquisition fails. It may signal a semaphore when it is not held. Should be backported to all supported versions. Closes scylladb/scylladb#19699 * github.com:scylladb/scylladb: test: add test to check that coordinator lwt semaphore continues functioning after locking failures paxos: do not signal semaphore if it was not acquired	2024-07-18 14:58:31 +03:00
Dawid Medrek	8b6e887e02	db/hints: Verify that Scylla limits the concurrency of written hints In `6e79d64`, the behavior of `manager::too_many_in_flight_hints_for()` was accidentally modified. It remained unnoticed for some time and then fixed. In this commit, we add a test verifying that the concurrency of hints being written to disk is indeed limited and the limitations are imposed properly.	2024-07-18 13:49:29 +02:00
Kefu Chai	db56af2e41	replication_strategy: mark fmt::formatter<..>::format() const since fmt 11, it is required that the format() to be const, otherwise its caller in fmt library would not be able to call it. and compile would fail like: ``` /home/kefu/.local/bin/clang++ -DFMT_SHARED -DSCYLLA_BUILD_MODE=release -DSEASTAR_API_LEVEL=7 -DSEASTAR_LOGGER_COMPILE_TIME_FMT -DSEASTAR_LOGGER_TYPE_STDOUT -DSEASTAR_SCHEDULING_GROUPS_COUNT=16 -DSEASTAR_SSTRING -DXXH_PRIVATE_API -DCMAKE_INTDIR=\"RelWithDebInfo\" -I/home/kefu/dev/scylladb -I/home/kefu/dev/scylladb/seastar/include -I/home/kefu/dev/scylladb/build/seastar/gen/include -I/home/kefu/dev/scylladb/build/seastar/gen/src -I/home/kefu/dev/scylladb/build/gen -isystem /home/kefu/dev/scylladb/abseil -ffunction-sections -fdata-sections -O3 -g -gz -std=gnu++23 -fvisibility=hidden -Wall -Werror -Wextra -Wno-error=deprecated-declarations -Wimplicit-fallthrough -Wno-c++11-narrowing -Wno-deprecated-copy -Wno-mismatched-tags -Wno-missing-field-initializers -Wno-overloaded-virtual -Wno-unsupported-friend -Wno-enum-constexpr-conversion -Wno-unused-parameter -ffile-prefix-map=/home/kefu/dev/scylladb=. -march=westmere -Xclang -fexperimental-assignment-tracking=disabled -mllvm -inline-threshold=2500 -fno-slp-vectorize -U_FORTIFY_SOURCE -Werror=unused-result -MD -MT locator/CMakeFiles/scylla_locator.dir/RelWithDebInfo/abstract_replication_strategy.cc.o -MF locator/CMakeFiles/scylla_locator.dir/RelWithDebInfo/abstract_replication_strategy.cc.o.d -o locator/CMakeFiles/scylla_locator.dir/RelWithDebInfo/abstract_replication_strategy.cc.o -c /home/kefu/dev/scylladb/locator/abstract_replication_strategy.cc In file included from /home/kefu/dev/scylladb/locator/abstract_replication_strategy.cc:9: In file included from /home/kefu/dev/scylladb/locator/abstract_replication_strategy.hh:16: In file included from /home/kefu/dev/scylladb/gms/inet_address.hh:11: In file included from /usr/include/fmt/ostream.h:23: In file included from /usr/include/fmt/chrono.h:23: In file included from /usr/include/fmt/format.h:41: /usr/include/fmt/base.h:1393:23: error: no matching member function for call to 'format' 1393 \| ctx.advance_to(cf.format(static_cast<qualified_type>(arg), ctx)); \| ~~~^~~~~~ /usr/include/fmt/base.h:1374:21: note: in instantiation of function template specialization 'fmt::detail::value<fmt::context>::format_custom_arg<locator::vnode_effective_replication_map::factory_key, fmt::formatter<locator::vnode_effective_replication_map::factory_key>>' requested here 1374 \| custom.format = format_custom_arg< \| ^ /home/kefu/dev/scylladb/seastar/include/seastar/util/log.hh:299:33: note: in instantiation of function template specialization 'fmt::format_to<seastar::internal::log_buf::inserter_iterator &, locator::vnode_effective_replication_map::factory_key &, const void , 0>' requested here 299 \| return fmt::format_to(it, fmt.format, std::forward<Args>(args)...); \| ^ /home/kefu/dev/scylladb/seastar/include/seastar/util/log.hh:428:9: note: in instantiation of function template specialization 'seastar::logger::log<locator::vnode_effective_replication_map::factory_key &, const void >' requested here 428 \| log(log_level::debug, std::move(fmt), std::forward<Args>(args)...); \| ^ /home/kefu/dev/scylladb/locator/abstract_replication_strategy.cc:561:18: note: in instantiation of function template specialization 'seastar::logger::debug<locator::vnode_effective_replication_map::factory_key &, const void *>' requested here 561 \| rslogger.debug("create_effective_replication_map: found {} [{}]", key, fmt::ptr(erm.get())); \| ^ /home/kefu/dev/scylladb/locator/abstract_replication_strategy.hh:471:10: note: candidate function template not viable: 'this' argument has type 'const fmt::formatter<locator::vnode_effective_replication_map::factory_key>', but method is not marked const 471 \| auto format(const locator::vnode_effective_replication_map::factory_key& key, FormatContext& ctx) { \| ^ 1 error generated. ``` Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19768	2024-07-18 13:52:36 +03:00
Emil Maskovsky	a89facbc74	test: raft: fix the flaky `test_raft_recovery_stuck` Use the rolling restart to avoid spurious driver reconnects. This can be eventually reverted once the scylladb/python-driver#295 is fixed. Fixes scylladb/scylladb#19154	2024-07-17 09:16:06 +02:00
Emil Maskovsky	ef3393bd36	test: raft: code cleanup in `test_raft_recovery_stuck` Cleaning up the imports.	2024-07-17 09:09:46 +02:00
Lakshmi Narayanan Sreethar	7b58fa2534	sstables: use _origin in write path Now that the origin is available inside the sstable object, no need to pass it to the methods called in the write path. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:44:28 +05:30
Lakshmi Narayanan Sreethar	b762a09dcd	sstable::open_sstable: pass and store origin Pass origin when opening the sstable from the writer and store it in the sstable object. This will make the origin available for the entire write path. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:43:30 +05:30
Lakshmi Narayanan Sreethar	7d0f3ace4a	test/boost: test abort behaviour during index read Added a new boost test, index_reader_test, with a testcase to verifyi the abort behaviour during an index read using index_consume_entry_context. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:42:50 +05:30
Lakshmi Narayanan Sreethar	64dadd5ec2	sstables/index_reader: stop consuming index when abort has been requested When an abort is requested, stop further reading of the index file and throw and exception from index_consume_entry_context::process_state(). Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:42:50 +05:30
Lakshmi Narayanan Sreethar	c2524337a2	sstables::index_consume_entry_context: store abort_source Store abort source inside sstables::index_consume_entry_context, so that the next patch can implement cancelling the index read when abort is requested. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:42:50 +05:30
Lakshmi Narayanan Sreethar	587da62686	sstable: drop old filter only after the new filter is built during rebuild sstable::maybe_rebuild_filter_from_index drops the existing filter first and then rebuilds the new filter as the method is only called before the sstable is sealed. But to make the index read abortable, the old filter can be dropped only after the new filter is built so that in case if the index consumer gets aborted, we still have the old filter to write to disk. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:42:47 +05:30
Lakshmi Narayanan Sreethar	6a3e7a5e7a	sstables/sstables_manager: store abort_source in sstable_manager Add a new member that stores the abort_source. This can later be used by the sstables to check if an abort has been requested. Also implement sstables_manager::get_abort_source() that returns a const reference to the abort source. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:36:06 +05:30
Lakshmi Narayanan Sreethar	e2142974f8	replica/database: pass abort_source to database constructor This is in preparation for the following patch that adds abort_source variable to the sstables_manager. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:36:06 +05:30
Piotr Dulikowski	6af7882c59	db/view: drop view updates to replaced node marked as left When a node that is permanently down is replaced, it is marked as "left" but it still can be a replica of some tablets. We also don't keep IPs of nodes that have left and the `node` structure for such node returns an empty IP (all zeros) as the address. This interacts badly with the view update logic. The base replica paired with the left node might decide to generate a view update. Because storage proxy still uses IPs and not host IDs, it needs to obtain the view replica's IP and tell the storage proxy to write a view update to that node - so, it chooses 0.0.0.0. Apparently, storage proxy decides to write a hint towards this address - hinted handoff on the other hand operates on host IDs and not IPs, so it attempts to translate the IP back, which triggers an assertion as there is no replica with IP 0.0.0.0. As a quick workaround for this issue just drop view updates towards nodes which seem to have IPs that are all zeros. It would be more proper to keep the view updates as hints and replay them later to the new paired replica, but achieving this right now would require much more significant changes. For now, fixing a crash is more important than keeping views consistent with base replicas. Fixes: scylladb/scylladb#19439	2024-07-16 15:50:11 +02:00
Gleb Natapov	4178589826	test: add test to check that coordinator lwt semaphore continues functioning after locking failures	2024-07-16 12:32:25 +03:00
Gleb Natapov	87beebeed0	paxos: do not signal semaphore if it was not acquired The guard signals a semaphore during destruction if it is marked as locked, but currently it may be marked as locked even if locking failed. Fix this by using semaphore_units instead of managing the locked flag manually. Fixes: https://github.com/scylladb/scylladb/issues/19698	2024-07-16 12:32:25 +03:00
Kefu Chai	c911832ed9	github: do not run clang-tidy as a cron job we already run it for every pull request, so no need to run it periodically. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-07-15 19:19:49 +08:00
Kefu Chai	dc189c67a6	github: disable scheduled workflow on forks as these workflows are scheduled periodically, and if they fail, notifications are sent to the repo's owner. to minimize the surprises to the contributors using github, let's disable these workflows on fork repos. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-07-15 19:19:28 +08:00
Marcin Maliszkiewicz	395dec35c1	cql3: functions: replace template with std::function in with_udf_iter() Templates are slower to compile and more difficult to read, in this case generalization is not needed and can be replaced by std::function.	2024-07-15 09:39:20 +02:00
Marcin Maliszkiewicz	85d38e013c	cql3: functions: improve functions class constness handling Declares getters as const methods. Makes instance() function return const object so that it may only be modified via change_batch class.	2024-07-15 09:39:20 +02:00
Marcin Maliszkiewicz	b9861c0bb7	Revert "cql3: functions: make modification functions accessible only via batch class" This reverts commit `3f1c2fecc2`. This access control property will be implemented differently (by using const) in subsequent commit hence revert.	2024-07-15 09:39:20 +02:00
Dawid Medrek	7301a96ff4	db/hints: Coroutinize `hint_endpoint_manager::store_hint()`	2024-07-15 04:15:25 +02:00
Raphael S. Carvalho	8df7f78969	replica: rename for_each_const_compaction_group() use same name as non-const-qualified variant, by relying on overloading. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-07-12 16:33:34 -03:00
Raphael S. Carvalho	518677d7f9	replica: Fix comment about compaction group there's not a 1:1 relationship between compaction group count and tablet count. a tablet replica has a storage group instance, which may map to multiple compaction groups during split mode. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-07-12 16:24:51 -03:00
Raphael S. Carvalho	f139aa1df6	replica: remove unused compaction_group_vector Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-07-12 16:16:47 -03:00
Dawid Medrek	3e02e66ca8	db/hints: Move a constant value to the TU it's used in Until now, the constant `HINT_FILE_WRITE_TIMEOUT` was declared as a static member of `db::hints::manager`. However, the constant is only ever used in one translation unit, so it makes more sense to move it there and not include boilerplate in a header.	2024-07-12 13:08:33 +02:00
Calle Wilund	8295980d14	commitlog: Make max data lifetime runtime-configurable	2024-07-09 12:30:49 +00:00
Calle Wilund	0c6679e55f	db::config: Expose commitlog_max_data_lifetime_in_s parameter To allow user control of commitlog time based expiry. Set to 24h initially.	2024-07-09 12:30:48 +00:00
Calle Wilund	55d6afda6e	commitlog: Add optional max lifetime parameter to cl instance If set, any remaining segment that has data older than this threshold will request flushing, regardless of data pressure. I.e. even a system where nothing happends will after X seconds flush data to free up the commit log.	2024-07-09 12:30:48 +00:00

1475 changed files with 52613 additions and 18957 deletions

209

.clang-format Normal file

View File

@@ -0,0 +1,209 @@
 ---
 Language: Cpp
 AccessModifierOffset: -4
 AlignAfterOpenBracket: DontAlign
 AlignArrayOfStructures: None
 AlignConsecutiveAssignments:
   Enabled: false
   AcrossEmptyLines: false
   AcrossComments: false
   AlignCompound: false
   PadOperators: true
 AlignConsecutiveBitFields:
   Enabled: false
   AcrossEmptyLines: false
   AcrossComments: false
   AlignCompound: false
   PadOperators: false
 AlignConsecutiveDeclarations:
   Enabled: false
   AcrossEmptyLines: false
   AcrossComments: false
   AlignCompound: false
   PadOperators: false
 AlignConsecutiveMacros:
   Enabled: false
   AcrossEmptyLines: false
   AcrossComments: false
   AlignCompound: false
   PadOperators: false
 AlignConsecutiveShortCaseStatements:
   Enabled: false
   AcrossEmptyLines: false
   AcrossComments: false
   AlignCaseColons: false
 AlignEscapedNewlines: Right
 AlignOperands: Align
 AlignTrailingComments:
   Kind: Always
   OverEmptyLines: 0
 AllowAllArgumentsOnNextLine: true
 AllowAllParametersOfDeclarationOnNextLine: true
 AllowShortBlocksOnASingleLine: Never
 AllowShortCaseLabelsOnASingleLine: false
 AllowShortEnumsOnASingleLine: true
 AllowShortFunctionsOnASingleLine: None
 AllowShortIfStatementsOnASingleLine: Never
 AllowShortLambdasOnASingleLine: Empty
 AllowShortLoopsOnASingleLine: false
 AlwaysBreakAfterDefinitionReturnType: None
 AlwaysBreakAfterReturnType: None
 AlwaysBreakBeforeMultilineStrings: false
 AlwaysBreakTemplateDeclarations: Yes
 AttributeMacros:
   - __capability
 BinPackArguments: true
 BinPackParameters: true
 BitFieldColonSpacing: Both
 BraceWrapping:
   AfterCaseLabel: false
   AfterClass: false
   AfterControlStatement: Never
   AfterEnum: false
   AfterExternBlock: false
   AfterFunction: false
   AfterNamespace: false
   AfterObjCDeclaration: false
   AfterStruct: false
   AfterUnion: false
   BeforeCatch: false
   BeforeElse: false
   BeforeLambdaBody: false
   BeforeWhile: false
   IndentBraces: false
   SplitEmptyFunction: true
   SplitEmptyRecord: true
   SplitEmptyNamespace: true
 BreakAfterAttributes: Never
 BreakAfterJavaFieldAnnotations: false
 BreakArrays: true
 BreakBeforeBinaryOperators: None
 BreakBeforeConceptDeclarations: Always
 BreakBeforeBraces: Attach
 BreakBeforeInlineASMColon: OnlyMultiline
 BreakBeforeTernaryOperators: true
 BreakConstructorInitializers: BeforeComma
 BreakInheritanceList: BeforeColon
 BreakStringLiterals: true
 ColumnLimit: 160
 CommentPragmas: '^ IWYU pragma:'
 CompactNamespaces: false
 ConstructorInitializerIndentWidth: 4
 ContinuationIndentWidth: 8
 Cpp11BracedListStyle: true
 DerivePointerAlignment: false
 DisableFormat: false
 EmptyLineAfterAccessModifier: Never
 EmptyLineBeforeAccessModifier: LogicalBlock
 ExperimentalAutoDetectBinPacking: false
 FixNamespaceComments: true
 ForEachMacros:
   - foreach
   - Q_FOREACH
   - BOOST_FOREACH
 IfMacros:
   - KJ_IF_MAYBE
 IndentAccessModifiers: false
 IndentCaseBlocks: false
 IndentCaseLabels: false
 IndentExternBlock: AfterExternBlock
 IndentGotoLabels: true
 IndentPPDirectives: None
 IndentRequiresClause: true
 IndentWidth: 4
 IndentWrappedFunctionNames: false
 InsertBraces: false
 InsertNewlineAtEOF: true
 InsertTrailingCommas: None
 IntegerLiteralSeparator:
   Binary: 0
   BinaryMinDigits: 0
   Decimal: 0
   DecimalMinDigits: 0
   Hex: 0
   HexMinDigits: 0
 JavaScriptQuotes: Leave
 JavaScriptWrapImports: true
 KeepEmptyLinesAtTheStartOfBlocks: true
 KeepEmptyLinesAtEOF: false
 LambdaBodyIndentation: Signature
 LineEnding: DeriveLF
 MacroBlockBegin: ''
 MacroBlockEnd: ''
 MaxEmptyLinesToKeep: 2
 NamespaceIndentation: None
 PackConstructorInitializers: BinPack
 PenaltyBreakAssignment: 2
 PenaltyBreakBeforeFirstCallParameter: 19
 PenaltyBreakComment: 300
 PenaltyBreakFirstLessLess: 120
 PenaltyBreakOpenParenthesis: 0
 PenaltyBreakString: 1000
 PenaltyBreakTemplateDeclaration: 10
 PenaltyExcessCharacter: 1000000
 PenaltyIndentedWhitespace: 0
 PenaltyReturnTypeOnItsOwnLine: 60
 PointerAlignment: Left
 PPIndentWidth: -1
 QualifierAlignment: Leave
 ReferenceAlignment: Pointer
 ReflowComments: true
 RemoveBracesLLVM: false
 RemoveParentheses: Leave
 RemoveSemicolon: false
 RequiresClausePosition: OwnLine
 RequiresExpressionIndentation: OuterScope
 SeparateDefinitionBlocks: Leave
 ShortNamespaceLines: 1
 SortIncludes: Never
 SortJavaStaticImport: Before
 SortUsingDeclarations: Never
 SpaceAfterCStyleCast: false
 SpaceAfterLogicalNot: false
 SpaceAfterTemplateKeyword: true
 SpaceAroundPointerQualifiers: Default
 SpaceBeforeAssignmentOperators: true
 SpaceBeforeCaseColon: false
 SpaceBeforeCpp11BracedList: false
 SpaceBeforeCtorInitializerColon: true
 SpaceBeforeInheritanceColon: true
 SpaceBeforeJsonColon: false
 SpaceBeforeParens: ControlStatements
 SpaceBeforeParensOptions:
   AfterControlStatements: true
   AfterForeachMacros: true
   AfterFunctionDefinitionName: false
   AfterFunctionDeclarationName: false
   AfterIfMacros: true
   AfterOverloadedOperator: false
   AfterRequiresInClause: false
   AfterRequiresInExpression: false
   BeforeNonEmptyParentheses: false
 SpaceBeforeRangeBasedForLoopColon: true
 SpaceBeforeSquareBrackets: false
 SpaceInEmptyBlock: false
 SpacesBeforeTrailingComments: 1
 SpacesInAngles: Never
 SpacesInContainerLiterals: true
 SpacesInLineCommentPrefix:
   Minimum: 1
   Maximum: -1
 SpacesInParens: Never
 SpacesInParensOptions:
   InCStyleCasts: false
   InConditionalStatements: false
   InEmptyParentheses: false
   Other: false
 SpacesInSquareBrackets: false
 Standard: Latest
 TabWidth: 8
 UseTab: Never
 VerilogBreakBetweenInstancePorts: true
 WhitespaceSensitiveMacros:
   - BOOST_PP_STRINGIZE
   - CF_SWIFT_NAME
   - NS_SWIFT_NAME
   - PP_STRINGIZE
   - STRINGIZE
 ...

31

.github/CODEOWNERS vendored

View File

@@ -1,5 +1,5 @@
 # AUTH
 auth/* @elcallio @vladzcloudius
 auth/* @nuivall @ptrsmrn @KrzaQ
 # CACHE
 row_cache* @tgrabiec
@@ -7,9 +7,9 @@ row_cache* @tgrabiec
 test/boost/mvcc* @tgrabiec
 # CDC
 cdc/* @kbr- @elcallio @piodul @jul-stas
 test/cql/cdc_* @kbr- @elcallio @piodul @jul-stas
 test/boost/cdc_* @kbr- @elcallio @piodul @jul-stas
 cdc/* @kbr-scylla @elcallio @piodul
 test/cql/cdc_* @kbr-scylla @elcallio @piodul
 test/boost/cdc_* @kbr-scylla @elcallio @piodul
 # COMMITLOG / BATCHLOG
 db/commitlog/* @elcallio @eliransin
@@ -25,18 +25,18 @@ compaction/* @raphaelsc
 transport/*
 # CQL QUERY LANGUAGE
 cql3/* @tgrabiec
 cql3/* @tgrabiec @nuivall @ptrsmrn @KrzaQ
 # COUNTERS
 counters* @jul-stas
 tests/counter_test* @jul-stas
 counters* @nuivall @ptrsmrn @KrzaQ
 tests/counter_test* @nuivall @ptrsmrn @KrzaQ
 # DOCS
 docs/* @annastuchlik @tzach
 docs/alternator @annastuchlik @tzach @nyh @havaker @nuivall
 docs/alternator @annastuchlik @tzach @nyh @nuivall @ptrsmrn @KrzaQ
 # GOSSIP
 gms/* @tgrabiec @asias
 gms/* @tgrabiec @asias @kbr-scylla
 # DOCKER
 dist/docker/*
@@ -74,8 +74,8 @@ streaming/* @tgrabiec @asias
 service/storage_service.* @tgrabiec @asias
 # ALTERNATOR
 alternator/* @havaker @nuivall
 test/alternator/* @havaker @nuivall
 alternator/* @nyh @nuivall @ptrsmrn @KrzaQ
 test/alternator/* @nyh @nuivall @ptrsmrn @KrzaQ
 # HINTED HANDOFF
 db/hints/* @piodul @vladzcloudius @eliransin
@@ -91,11 +91,14 @@ test/boost/mutation_reader_test.cc @denesb
 test/boost/querier_cache_test.cc @denesb
 # PYTEST-BASED CQL TESTS
 test/cql-pytest/* @nyh
 test/cqlpy/* @nyh
 # RAFT
 raft/* @kbr- @gleb-cloudius @kostja
 test/raft/* @kbr- @gleb-cloudius @kostja
 raft/* @kbr-scylla @gleb-cloudius @kostja
 test/raft/* @kbr-scylla @gleb-cloudius @kostja
 # HEAT-WEIGHTED LOAD BALANCING
 db/heat_load_balance.* @nyh @gleb-cloudius
 # Tools
 tools/* @denesb

									
										15

.github/ISSUE_TEMPLATE.md
									
										vendored
									
												View File
											
				@@ -1,15 +0,0 @@

				This is Scylla's bug tracker, to be used for reporting bugs only.

				If you have a question about Scylla, and not a bug, please ask it in

				our mailing-list at scylladb-dev@googlegroups.com or in our slack channel.

				- [] I have read the disclaimer above, and I am reporting a suspected malfunction in Scylla.

				*Installation details*

				Scylla version (or git commit hash):

				Cluster size:

				OS (RHEL/CentOS/Ubuntu/AWS AMI):

				*Hardware details (for performance issues)*          Delete if unneeded

				Platform (physical/VM/cloud instance type/docker):

				Hardware: sockets= cores= hyperthreading= memory=

				Disks: (SSD/HDD, count)

									
										86

.github/ISSUE_TEMPLATE/bug_report.yml
									
										vendored
									
										Normal file
									
												View File
												
				@@ -0,0 +1,86 @@

				name: "Report a bug"

				description: "File a bug report."

				title: "[Bug]: "

				type: "bug"

				labels: bug

				body:

				  - type: checkboxes

				    id: terms

				    attributes:

				      label: Code of Conduct

				      description: "This is Scylla's bug tracker, to be used for reporting bugs only.

				If you have a question about Scylla, and not a bug, please ask it in

				our forum at https://forum.scylladb.com/ or in our slack channel https://slack.scylladb.com/ "

				      options:

				        - label: I have read the disclaimer above and am reporting a suspected malfunction in Scylla.

				          required: true

				  - type: input

				    id: product-version

				    attributes:

				      label: product version

				      description: Scylla version (or git commit hash)

				      placeholder: ex. scylla-6.1.1

				    validations:

				      required: true

				  - type: input

				    id: cluster-size

				    attributes:

				      label: Cluster Size

				    validations:

				      required: true  

				  - type: input

				    id: os

				    attributes:

				      label: OS

				      placeholder: RHEL/CentOS/Ubuntu/AWS AMI

				    validations:

				      required: true

				  - type: textarea

				    id: additional-data

				    attributes:

				      label: Additional Environmental Data

				      #description: 

				      placeholder: Add additional data

				      value: "Platform (physical/VM/cloud instance type/docker):\n

				Hardware: sockets=   cores=   hyperthreading=   memory=\n

				Disks: (SSD/HDD, count)"

				    validations:

				      required: false

				  - type: textarea

				    id: reproducer-steps

				    attributes:

				      label: Reproduction Steps

				      placeholder: Describe how to reproduce the problem

				      value: "The steps to reproduce the problem are:"

				    validations:

				      required: true

				  - type: textarea

				    id: the-problem

				    attributes:

				      label: What is the problem?

				      placeholder: Describe the problem you found

				      value: "The problem is that"

				    validations:

				      required: true

				  - type: textarea

				    id: what-happened

				    attributes:

				      label: Expected behavior?

				      placeholder: Describe what should have happened

				      value: "I expected that "

				    validations:

				      required: true

				  - type: textarea

				    id: logs

				    attributes:

				      label: Relevant log output

				      description: Please copy and paste any relevant log output. This will be automatically formatted into code, so no need for backticks.

				      render: shell

									
										9

.github/dependabot.yml
									
										vendored
									
										Normal file
									
												View File
												
				@@ -0,0 +1,9 @@

				version: 2

				updates:

				- package-ecosystem: "pip"

				  directory: "/docs"

				  schedule:

				    interval: "daily"

				  allow:

				  - dependency-name: "sphinx-scylladb-theme"

				  - dependency-name: "sphinx-multiversion-scylla"

									
										58

.github/mergify.yml
									
										vendored
									
												View File
												
				@@ -15,7 +15,7 @@ pull_request_rules:

				        - closed

				    actions:

				      delete_head_branch:

				  - name: Automate backport pull request 5.2

				  - name: Automate backport pull request 6.2

				    conditions:

				      - or:

				        - closed

				@@ -23,36 +23,11 @@ pull_request_rules:

				      - or:

				          - base=master

				          - base=next

				      - label=backport/5.2 # The PR must have this label to trigger the backport

				      - label=backport/6.2 # The PR must have this label to trigger the backport

				      - label=promoted-to-master

				    actions:

				      copy:

				        title: "[Backport 5.2] {{ title }}"

				        body: |

				          {{ body }}

				          {% for c in commits %}

				          (cherry picked from commit {{ c.sha }})

				          {% endfor %}

				           Refs #{{number}}

				        branches:

				          - branch-5.2

				        assignees:

				          - "{{ author }}"

				  - name: Automate backport pull request 5.4

				    conditions:

				      - or:

				        - closed

				        - merged

				      - or:

				          - base=master

				          - base=next

				      - label=backport/5.4 # The PR must have this label to trigger the backport

				      - label=promoted-to-master

				    actions:

				      copy:

				        title: "[Backport 5.4] {{ title }}"

				        title: "[Backport 6.2] {{ title }}"

				        body: |

				          {{ body }}

				@@ -62,7 +37,32 @@ pull_request_rules:

				          Refs #{{number}}

				        branches:

				          - branch-5.4

				          - branch-6.2

				        assignees:

				          - "{{ author }}"

				  - name: Automate backport pull request 6.1

				    conditions:

				      - or:

				        - closed

				        - merged

				      - or:

				          - base=master

				          - base=next

				      - label=backport/6.1 # The PR must have this label to trigger the backport

				      - label=promoted-to-master

				    actions:

				      copy:

				        title: "[Backport 6.1] {{ title }}"

				        body: |

				          {{ body }}

				          {% for c in commits %}

				          (cherry picked from commit {{ c.sha }})

				          {% endfor %}

				           Refs #{{number}}

				        branches:

				          - branch-6.1

				        assignees:

				          - "{{ author }}"

				  - name: Automate backport pull request 6.0

									
										181

.github/scripts/auto-backport.py
									
										vendored
									
										Executable file
									
												View File
												
				@@ -0,0 +1,181 @@

				#!/usr/bin/env python3

				import argparse

				import os

				import re

				import sys

				import tempfile

				import logging

				from github import Github, GithubException

				from git import Repo, GitCommandError

				logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')

				try:

				    github_token = os.environ["GITHUB_TOKEN"]

				except KeyError:

				    print("Please set the 'GITHUB_TOKEN' environment variable")

				    sys.exit(1)

				def is_pull_request():

				    return '--pull-request' in sys.argv[1:]

				def parse_args():

				    parser = argparse.ArgumentParser()

				    parser.add_argument('--repo', type=str, required=True, help='Github repository name')

				    parser.add_argument('--base-branch', type=str, default='refs/heads/master', help='Base branch')

				    parser.add_argument('--commits', default=None, type=str, help='Range of promoted commits.')

				    parser.add_argument('--pull-request', type=int, help='Pull request number to be backported')

				    parser.add_argument('--head-commit', type=str, required=is_pull_request(), help='The HEAD of target branch after the pull request specified by --pull-request is merged')

				    return parser.parse_args()

				def create_pull_request(repo, new_branch_name, base_branch_name, pr, backport_pr_title, commits, is_draft=False):

				    pr_body = f'{pr.body}\n\n'

				    for commit in commits:

				        pr_body += f'- (cherry picked from commit {commit})\n\n'

				    pr_body += f'Parent PR: #{pr.number}'

				    try:

				        backport_pr = repo.create_pull(

				            title=backport_pr_title,

				            body=pr_body,

				            head=f'scylladbbot:{new_branch_name}',

				            base=base_branch_name,

				            draft=is_draft

				        )

				        logging.info(f"Pull request created: {backport_pr.html_url}")

				        backport_pr.add_to_assignees(pr.user)

				        logging.info(f"Assigned PR to original author: {pr.user}")

				        return backport_pr

				    except GithubException as e:

				        if 'A pull request already exists' in str(e):

				            logging.warning(f'A pull request already exists for {pr.user}:{new_branch_name}')

				        else:

				            logging.error(f'Failed to create PR: {e}')

				def get_pr_commits(repo, pr, stable_branch, start_commit=None):

				    commits = []

				    if pr.merged:

				        merge_commit = repo.get_commit(pr.merge_commit_sha)

				        if len(merge_commit.parents) > 1:  # Check if this merge commit includes multiple commits

				            commits.append(pr.merge_commit_sha)

				        else:

				            if start_commit:

				                promoted_commits = repo.compare(start_commit, stable_branch).commits

				            else:

				                promoted_commits = repo.get_commits(sha=stable_branch)

				            for commit in pr.get_commits():

				                for promoted_commit in promoted_commits:

				                    commit_title = commit.commit.message.splitlines()[0]

				                    # In Scylla-pkg and scylla-dtest, for example,

				                    # we don't create a merge commit for a PR with multiple commits,

				                    # according to the GitHub API, the last commit will be the merge commit,

				                    # which is not what we need when backporting (we need all the commits).

				                    # So here, we are validating the correct SHA for each commit so we can cherry-pick

				                    if promoted_commit.commit.message.startswith(commit_title):

				                        commits.append(promoted_commit.sha)

				    elif pr.state == 'closed':

				        events = pr.get_issue_events()

				        for event in events:

				            if event.event == 'closed':

				                commits.append(event.commit_id)

				    return commits

				def create_pr_comment_and_remove_label(pr, comment_body):

				    labels = pr.get_labels()

				    pattern = re.compile(r"backport/\d+\.\d+$")

				    for label in labels:

				        if pattern.match(label.name):

				            print(f"Removing label: {label.name}")

				            comment_body += f'- {label.name}\n'

				            pr.remove_from_labels(label)

				    pr.create_issue_comment(comment_body)

				def backport(repo, pr, version, commits, backport_base_branch):

				    new_branch_name = f'backport/{pr.number}/to-{version}'

				    backport_pr_title = f'[Backport {version}] {pr.title}'

				    repo_url = f'https://scylladbbot:{github_token}@github.com/{repo.full_name}.git'

				    fork_repo = f'https://scylladbbot:{github_token}@github.com/scylladbbot/{repo.name}.git'

				    with (tempfile.TemporaryDirectory() as local_repo_path):

				        try:

				            repo_local = Repo.clone_from(repo_url, local_repo_path, branch=backport_base_branch)

				            repo_local.git.checkout(b=new_branch_name)

				            is_draft = False

				            for commit in commits:

				                try:

				                    repo_local.git.cherry_pick(commit, '-m1', '-x')

				                except GitCommandError as e:

				                    logging.warning(f'Cherry-pick conflict on commit {commit}: {e}')

				                    is_draft = True

				                    repo_local.git.add(A=True)

				                    repo_local.git.cherry_pick('--continue')

				            if not repo.private and not repo.has_in_collaborators(pr.user.login):

				                repo.add_to_collaborators(pr.user.login, permission="push")

				                comment = f':warning:  @{pr.user.login} you have been added as collaborator to scylladbbot fork '

				                comment += f'Please check your inbox and approve the invitation, once it is done, please add the backport labels again'

				                create_pr_comment_and_remove_label(pr, comment)

				                return

				            repo_local.git.push(fork_repo, new_branch_name, force=True)

				            create_pull_request(repo, new_branch_name, backport_base_branch, pr, backport_pr_title, commits,

				                                is_draft=is_draft)

				        except GitCommandError as e:

				            logging.warning(f"GitCommandError: {e}")

				def main():

				    args = parse_args()

				    base_branch = args.base_branch.split('/')[2]

				    promoted_label = 'promoted-to-master'

				    repo_name = args.repo

				    if 'scylla-enterprise' in args.repo:

				        promoted_label = 'promoted-to-enterprise'

				    stable_branch = base_branch

				    backport_branch = 'branch-'

				    backport_label_pattern = re.compile(r'backport/\d+\.\d+$')

				    g = Github(github_token)

				    repo = g.get_repo(repo_name)

				    closed_prs = []

				    start_commit = None

				    if args.commits:

				        start_commit, end_commit = args.commits.split('..')

				        commits = repo.compare(start_commit, end_commit).commits

				        for commit in commits:

				            match = re.search(rf"Closes .*#([0-9]+)", commit.commit.message, re.IGNORECASE)

				            if match:

				                pr_number = int(match.group(1))

				                pr = repo.get_pull(pr_number)

				                closed_prs.append(pr)

				    if args.pull_request:

				        start_commit = args.head_commit

				        pr = repo.get_pull(args.pull_request)

				        closed_prs = [pr]

				    for pr in closed_prs:

				        labels = [label.name for label in pr.labels]

				        backport_labels = [label for label in labels if backport_label_pattern.match(label)]

				        if promoted_label not in labels:

				            print(f'no {promoted_label} label: {pr.number}')

				            continue

				        if not backport_labels:

				            print(f'no backport label: {pr.number}')

				            continue

				        commits = get_pr_commits(repo, pr, stable_branch, start_commit)

				        logging.info(f"Found PR #{pr.number} with commit {commits} and the following labels: {backport_labels}")

				        for backport_label in backport_labels:

				            version = backport_label.replace('backport/', '')

				            backport_base_branch = backport_label.replace('backport/', backport_branch)

				            backport(repo, pr, version, commits, backport_base_branch)

				if __name__ == "__main__":

				    main()

									
										23

.github/scripts/label_promoted_commits.py
									
										vendored
									
												View File
												
				@@ -16,13 +16,8 @@ def parser():

				    parser = argparse.ArgumentParser()

				    parser.add_argument('--repository', type=str, required=True,

				                        help='Github repository name (e.g., scylladb/scylladb)')

				    parser.add_argument('--commit_before_merge', type=str, required=True, help='Git commit ID to start labeling from ('

				                                                                               'newest commit).')

				    parser.add_argument('--commit_after_merge', type=str, required=True,

				                        help='Git commit ID to end labeling at (oldest '

				                             'commit, exclusive).')

				    parser.add_argument('--update_issue', type=bool, default=False, help='Set True to update issues when backport was '

				                                                                         'done')

				    parser.add_argument('--commits', type=str, required=True, help='Range of promoted commits.')

				    parser.add_argument('--label', type=str, default='promoted-to-master', help='Label to use')

				    parser.add_argument('--ref', type=str, required=True, help='PR target branch')

				    return parser.parse_args()

				@@ -53,12 +48,14 @@ def main():

				    target_branch = re.search(r'branch-(\d+\.\d+)', args.ref)

				    g = Github(github_token)

				    repo = g.get_repo(args.repository, lazy=False)

				    commits = repo.compare(head=args.commit_after_merge, base=args.commit_before_merge)

				    start_commit, end_commit = args.commits.split('..')

				    commits = repo.compare(start_commit, end_commit).commits

				    processed_prs = set()

				    # Print commit information

				    for commit in commits.commits:

				    for commit in commits:

				        print(f'Commit sha is: {commit.sha}')

				        match = pr_pattern.search(commit.commit.message)

				        pr_last_line = commit.commit.message.splitlines()[-1]

				        match = pr_pattern.search(pr_last_line)

				        if match:

				            pr_number = int(match.group(1))

				            if pr_number in processed_prs:

				@@ -66,13 +63,13 @@ def main():

				            if target_branch:

				                pr = repo.get_pull(pr_number)

				                branch_name = target_branch[1]

				                refs_pr = re.findall(r'Refs (?:#|https.*?)(\d+)', pr.body)

				                refs_pr = re.findall(r'Parent PR: (?:#|https.*?)(\d+)', pr.body)

				                if refs_pr:

				                    print(f'branch-{target_branch.group(1)}, pr number is: {pr_number}')

				                    # 1. change the backport label of the parent PR to note that

				                    #    we've merge the corresponding backport PR

				                    #    we've merged the corresponding backport PR

				                    # 2. close the backport PR and leave a comment on it to note

				                    #    that it has been merged with a certain git commit,

				                    #    that it has been merged with a certain git commit.

				                    ref_pr_number = refs_pr[0]

				                    mark_backport_done(repo, ref_pr_number, branch_name)

				                    comment = f'Closed via {commit.sha}'

									
										51

.github/workflows/add-label-when-promoted.yaml
									
										vendored
									
												View File
												
				@@ -5,9 +5,10 @@ on:

				    branches:

				      - master

				      - branch-*.*

				env:

				  DEFAULT_BRANCH: 'master'

				      - enterprise

				    pull_request_target:

				      types: [labeled]

				      branches: [master, next, enterprise]

				jobs:

				  check-commit:

				@@ -20,17 +21,51 @@ jobs:

				        env:

				          GITHUB_CONTEXT: ${{ toJson(github) }}

				        run: echo "$GITHUB_CONTEXT"

				      - name: Set Default Branch

				        id: set_branch

				        run: |

				          if [[ "${{ github.repository }}" == *enterprise* ]]; then

				            echo "DEFAULT_BRANCH=enterprise" >> $GITHUB_ENV

				          else

				            echo "DEFAULT_BRANCH=master" >> $GITHUB_ENV

				          fi

				      - name: Checkout repository

				        uses: actions/checkout@v4

				        with:

				          repository: ${{ github.repository }}

				          ref: ${{ env.DEFAULT_BRANCH }}

				          token: ${{ secrets.AUTO_BACKPORT_TOKEN }}

				          fetch-depth: 0  # Fetch all history for all tags and branches

				      - name: Set up Git identity

				        run: |

				          git config --global user.name "GitHub Action"

				          git config --global user.email "action@github.com"

				          git config --global merge.conflictstyle diff3

				      - name: Install dependencies

				        run: sudo apt-get install -y python3-github

				        run: sudo apt-get install -y python3-github python3-git

				      - name: Run python script

				        if: github.event_name == 'push'

				        env:

				          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}

				        run: python .github/scripts/label_promoted_commits.py --commit_before_merge ${{ github.event.before }} --commit_after_merge ${{ github.event.after }} --repository ${{ github.repository }} --ref ${{ github.ref }}

				          GITHUB_TOKEN: ${{ secrets.AUTO_BACKPORT_TOKEN }}

				        run: python .github/scripts/label_promoted_commits.py  --commits ${{ github.event.before }}..${{ github.sha }} --repository ${{ github.repository }} --ref ${{ github.ref }}

				      - name: Run auto-backport.py when promotion completed

				        if: github.event_name == 'push' && github.ref == 'refs/heads/${{ env.DEFAULT_BRANCH }}'

				        env:

				          GITHUB_TOKEN: ${{ secrets.AUTO_BACKPORT_TOKEN }}

				        run: python .github/scripts/auto-backport.py --repo ${{ github.repository }} --base-branch ${{ github.ref }} --commits ${{ github.event.before }}..${{ github.sha }}

				      - name: Check if label starts with 'backport/' and contains digits

				        id: check_label

				        run: |

				          label_name="${{ github.event.label.name }}"

				          if [[ "$label_name" =~ ^backport/[0-9]+\.[0-9]+$ ]]; then

				            echo "Label matches backport/X.X pattern."

				            echo "backport_label=true" >> $GITHUB_OUTPUT

				          else

				            echo "Label does not match the required pattern."

				            echo "backport_label=false" >> $GITHUB_OUTPUT

				          fi

				      - name: Run auto-backport.py when label was added

				        if: github.event_name == 'pull_request_target' && steps.check_label.outputs.backport_label == 'true' && (github.event.pull_request.state == 'closed' && github.event.pull_request.merged == true)

				        env:

				          GITHUB_TOKEN: ${{ secrets.AUTO_BACKPORT_TOKEN }}

				        run: python .github/scripts/auto-backport.py --repo ${{ github.repository }} --base-branch ${{ github.ref }} --pull-request ${{ github.event.pull_request.number }} --head-commit ${{ github.event.pull_request.base.sha }}

									
										9

.github/workflows/backport-pr-fixes-validation.yaml
									
										vendored
									
												View File
												
				@@ -22,5 +22,12 @@ jobs:

				            const regex = new RegExp(pattern);

				            if (!regex.test(body)) {

				              core.setFailed("PR body does not contain a valid 'Fixes' reference.");

				              const error = "PR body does not contain a valid 'Fixes' reference.";

				              core.setFailed(error);

				              await github.rest.issues.createComment({

				                issue_number: context.issue.number,

				                owner: context.repo.owner,

				                repo: context.repo.repo,

				                body: `:warning: ${error}`

				              });

				            }

									
										8

.github/workflows/build-scylla.yaml
									
										vendored
									
												View File
												
				@@ -13,10 +13,14 @@ on:

				        value: ${{ jobs.build.outputs.md5sum }}

				jobs:

				  read-toolchain:

				    uses: ./.github/workflows/read-toolchain.yaml

				  build:

				    if: github.repository == 'scylladb/scylladb'

				    needs:

				      - read-toolchain

				    runs-on: ubuntu-latest

				    # be consistent with tools/toolchain/image

				    container: scylladb/scylla-toolchain:fedora-40-20240621

				    container: ${{ needs.read-toolchain.outputs.image }}

				    outputs:

				      md5sum: ${{ steps.checksum.outputs.md5sum }}

				    steps:

									
										3

.github/workflows/clang-nightly.yaml
									
										vendored
									
												View File
												
				@@ -7,7 +7,7 @@ on:

				env:

				  # use the development branch explicitly

				  CLANG_VERSION: 19

				  CLANG_VERSION: 20

				  BUILD_DIR: build

				permissions: {}

				@@ -20,6 +20,7 @@ concurrency:

				jobs:

				  clang-dev:

				    name: Build with clang nightly

				    if: github.repository == 'scylladb/scylladb'

				    runs-on: ubuntu-latest

				    container: fedora:40

				    strategy:

									
										7

.github/workflows/clang-tidy.yaml
									
										vendored
									
												View File
												
				@@ -10,9 +10,9 @@ on:

				      - 'docs/**'

				      - '.github/**'

				  workflow_dispatch:

				  schedule:

				    # only at 5AM Saturday

				    - cron: '0 5 * * SAT'

				  issue_comment:

				    types:

				      - created

				env:

				  BUILD_TYPE: RelWithDebInfo

				@@ -28,6 +28,7 @@ concurrency:

				jobs:

				  read-toolchain:

				    if: github.event_name == 'pull_request' || (github.event.issue.pull_request && startsWith(github.event.comment.body, '/clang-tidy'))

				    uses: ./.github/workflows/read-toolchain.yaml

				  clang-tidy:

				    name: Run clang-tidy

									
										45

.github/workflows/conflict_reminder.yaml
									
										vendored
									
										Normal file
									
												View File
												
				@@ -0,0 +1,45 @@

				name: Notify PR Authors of Conflicts

				on:

				  schedule:

				    - cron: '0 10 * * 1,4'  # Runs every Monday and Thursday at 10:00am

				  workflow_dispatch:      # Manual trigger for testing

				jobs:

				  notify_conflict_prs:

				    runs-on: ubuntu-latest

				    steps:

				      - name: Notify PR Authors of Conflicts

				        uses: actions/github-script@v7

				        with:

				          script: |

				            const prs = await github.paginate(github.rest.pulls.list, {

				              owner: context.repo.owner,

				              repo: context.repo.repo,

				              state: 'open',

				              per_page: 100

				            });

				            const branchPrefix = 'branch-';

				            const threeDaysAgo = new Date();

				            const conflictLabel = 'conflicts';          

				            threeDaysAgo.setDate(threeDaysAgo.getDate() - 3);

				            for (const pr of prs) {

				              if (!pr.base.ref.startsWith(branchPrefix)) continue;

				              const hasConflictLabel = pr.labels.some(label => label.name === conflictLabel);

				              if (!hasConflictLabel) continue;

				              const updatedDate = new Date(pr.updated_at);

				              if (updatedDate >= threeDaysAgo) continue;

				              if (pr.assignee === null) continue;

				              const assignee = pr.assignee;

				              if (assignee) {

				                await github.rest.issues.createComment({

				                  owner: context.repo.owner,

				                  repo: context.repo.repo,

				                  issue_number: pr.number,

				                  body: `@${assignee}, this PR has been open with conflicts. Please resolve the conflicts so we can merge it.`,

				                });

				                console.log(`Notified @${assignee} for PR #${pr.number}`);

				              } 

				            }

				            console.log(`Total PRs checked: ${prs.length}`);

									
										3

.github/workflows/docs-pr.yaml
									
										vendored
									
												View File
												
				@@ -12,7 +12,8 @@ on:

				      - enterprise

				    paths:

				      - "docs/**"

				      - "db/config.hh"

				      - "db/config.cc"

				jobs:

				  build:

				    runs-on: ubuntu-latest

									
										4

.github/workflows/iwyu.yaml
									
										vendored
									
												View File
												
				@@ -9,7 +9,9 @@ env:

				  BUILD_TYPE: RelWithDebInfo

				  BUILD_DIR: build

				  CLEANER_OUTPUT_PATH: build/clang-include-cleaner.log

				  CLEANER_DIRS: test/unit exceptions alternator api auth cdc compaction

				  # the "idl" subdirectory does not contain C++ source code. the .hh files in it are

				  # supposed to be processed by idl-compiler.py, so we don't check them using the cleaner

				  CLEANER_DIRS: test/unit exceptions alternator api auth cdc compaction db dht gms index lang

				permissions: {}

									
										1

.github/workflows/reproducible-build.yaml
									
										vendored
									
												View File
												
				@@ -19,6 +19,7 @@ jobs:

				    with:

				      build_mode: release

				  compare-checksum:

				    if: github.repository == 'scylladb/scylladb'

				    runs-on: ubuntu-latest

				    needs:

				      - build-a

5

.gitignore vendored

View File

@@ -3,6 +3,7 @@
 .settings
 build
 build.ninja
 cmake-build-*
 build.ninja.new
 cscope.*
 /debian/
@@ -13,13 +14,14 @@ dist/ami/scylla_deploy.sh
 Cql.tokens
 .kdev4
 *.kdev4
 .idea
 CMakeLists.txt.user
 .cache
 .tox
 *.egg-info
 __pycache__CMakeLists.txt.user
 .gdbinit
 resources
 /resources
 .pytest_cache
 /expressions.tokens
 tags
@@ -32,3 +34,4 @@ compile_commands.json
 .mypy_cache
 .envrc
 clang_build
 .idea/

3

.gitmodules vendored

View File

@@ -9,9 +9,6 @@
 [submodule "abseil"]
 	path = abseil
 	url = ../abseil-cpp
 [submodule "scylla-jmx"]
 	path = tools/jmx
 	url = ../scylla-jmx
 [submodule "scylla-tools"]
 	path = tools/java
 	url = ../scylla-tools-java

									
										92

CMakeLists.txt
									
												View File
												
				@@ -2,8 +2,6 @@ cmake_minimum_required(VERSION 3.27)

				project(scylla)

				include(CTest)

				list(APPEND CMAKE_MODULE_PATH

				  ${CMAKE_CURRENT_SOURCE_DIR}/cmake

				  ${CMAKE_CURRENT_SOURCE_DIR}/seastar/cmake)

				@@ -25,7 +23,8 @@ if(DEFINED CMAKE_BUILD_TYPE)

				endif(DEFINED CMAKE_BUILD_TYPE)

				include(mode.common)

				if(CMAKE_CONFIGURATION_TYPES)

				get_property(is_multi_config GLOBAL PROPERTY GENERATOR_IS_MULTI_CONFIG)

				if(is_multi_config)

				    foreach(config ${CMAKE_CONFIGURATION_TYPES})

				        include(mode.${config})

				        list(APPEND scylla_build_modes ${scylla_build_mode_${config}})

				@@ -49,26 +48,70 @@ set(CMAKE_CXX_EXTENSIONS ON CACHE INTERNAL "")

				set(CMAKE_CXX_SCAN_FOR_MODULES OFF CACHE INTERNAL "")

				set(CMAKE_CXX_VISIBILITY_PRESET hidden)

				set(Seastar_TESTING ON CACHE BOOL "" FORCE)

				set(Seastar_API_LEVEL 7 CACHE STRING "" FORCE)

				set(Seastar_DEPRECATED_OSTREAM_FORMATTERS OFF CACHE BOOL "" FORCE)

				set(Seastar_APPS ON CACHE BOOL "" FORCE)

				set(Seastar_EXCLUDE_APPS_FROM_ALL ON CACHE BOOL "" FORCE)

				set(Seastar_EXCLUDE_TESTS_FROM_ALL ON CACHE BOOL "" FORCE)

				set(Seastar_UNUSED_RESULT_ERROR ON CACHE BOOL "" FORCE)

				add_subdirectory(seastar)

				if(is_multi_config)

				    find_package(Seastar)

				    # this is atypical compared to standard ExternalProject usage:

				    # - Seastar's build system should already be configured at this point.

				    # - We maintain separate project variants for each configuration type.

				    #

				    # Benefits of this approach:

				    # - Allows the parent project to consume the compile options exposed by

				    #   .pc file. as the compile options vary from one config to another.

				    # - Allows application of config-specific settings

				    # - Enables building Seastar within the parent project's build system

				    # - Facilitates linking of artifacts with the external project target,

				    #   establishing proper dependencies between them

				    include(ExternalProject)

				    ExternalProject_Add(Seastar

				        SOURCE_DIR "${PROJECT_SOURCE_DIR}/seastar"

				        BINARY_DIR "${CMAKE_BINARY_DIR}/$<CONFIG>/seastar"

				        CONFIGURE_COMMAND ""

				        BUILD_COMMAND ${CMAKE_COMMAND} --build <BINARY_DIR>

				          --target seastar

				          --target seastar_testing

				          --target seastar_perf_testing

				          --target app_iotune

				        BUILD_ALWAYS ON

				        BUILD_BYPRODUCTS

				          <BINARY_DIR>/libseastar.$<IF:$<CONFIG:Debug,Dev>,so,a>

				          <BINARY_DIR>/libseastar_testing.$<IF:$<CONFIG:Debug,Dev>,so,a>

				          <BINARY_DIR>/libseastar_perf_testing.$<IF:$<CONFIG:Debug,Dev>,so,a>

				          <BINARY_DIR>/apps/iotune/iotune

				          <BINARY_DIR>/gen/include/seastar/http/chunk_parsers.hh

				          <BINARY_DIR>/gen/include/seastar/http/request_parser.hh

				          <BINARY_DIR>/gen/include/seastar/http/response_parser.hh

				        INSTALL_COMMAND "")

				    add_dependencies(Seastar::seastar Seastar)

				    add_dependencies(Seastar::seastar_testing Seastar)

				else()

				    set(Seastar_TESTING ON CACHE BOOL "" FORCE)

				    set(Seastar_API_LEVEL 7 CACHE STRING "" FORCE)

				    set(Seastar_DEPRECATED_OSTREAM_FORMATTERS OFF CACHE BOOL "" FORCE)

				    set(Seastar_APPS ON CACHE BOOL "" FORCE)

				    set(Seastar_EXCLUDE_APPS_FROM_ALL ON CACHE BOOL "" FORCE)

				    set(Seastar_EXCLUDE_TESTS_FROM_ALL ON CACHE BOOL "" FORCE)

				    set(Seastar_IO_URING OFF CACHE BOOL "" FORCE)

				    set(Seastar_SCHEDULING_GROUPS_COUNT 16 CACHE STRING "" FORCE)

				    set(Seastar_UNUSED_RESULT_ERROR ON CACHE BOOL "" FORCE)

				    add_subdirectory(seastar)

				    target_compile_definitions (seastar

				      PRIVATE

				        SEASTAR_NO_EXCEPTION_HACK)

				endif()

				set(ABSL_PROPAGATE_CXX_STD ON CACHE BOOL "" FORCE)

				find_package(Sanitizers QUIET)

				set(sanitizer_cxx_flags

				    $<$<IN_LIST:$<CONFIG>,Debug;Sanitize>:$<TARGET_PROPERTY:Sanitizers::address,INTERFACE_COMPILE_OPTIONS>;$<TARGET_PROPERTY:Sanitizers::undefined_behavior,INTERFACE_COMPILE_OPTIONS>>)

				    $<$<CONFIG:Debug,Sanitize>:$<TARGET_PROPERTY:Sanitizers::address,INTERFACE_COMPILE_OPTIONS>;$<TARGET_PROPERTY:Sanitizers::undefined_behavior,INTERFACE_COMPILE_OPTIONS>>)

				if(CMAKE_CXX_COMPILER_ID STREQUAL "GNU")

				    set(ABSL_GCC_FLAGS ${sanitizer_cxx_flags})

				elseif(CMAKE_CXX_COMPILER_ID STREQUAL "Clang")

				    set(ABSL_LLVM_FLAGS ${sanitizer_cxx_flags})

				endif()

				set(ABSL_DEFAULT_LINKOPTS

				    $<$<IN_LIST:$<CONFIG>,Debug;Sanitize>:$<TARGET_PROPERTY:Sanitizers::address,INTERFACE_LINK_LIBRARIES>;$<TARGET_PROPERTY:Sanitizers::undefined_behavior,INTERFACE_LINK_LIBRARIES>>)

				    $<$<CONFIG:Debug,Sanitize>:$<TARGET_PROPERTY:Sanitizers::address,INTERFACE_LINK_LIBRARIES>;$<TARGET_PROPERTY:Sanitizers::undefined_behavior,INTERFACE_LINK_LIBRARIES>>)

				add_subdirectory(abseil)

				add_library(absl-headers INTERFACE)

				target_include_directories(absl-headers SYSTEM INTERFACE

				@@ -95,12 +138,13 @@ target_link_libraries(Boost::regex

				find_package(Lua REQUIRED)

				find_package(ZLIB REQUIRED)

				find_package(ICU COMPONENTS uc i18n REQUIRED)

				find_package(fmt 9.0.0 REQUIRED)

				find_package(fmt 10.0.0 REQUIRED)

				find_package(libdeflate REQUIRED)

				find_package(libxcrypt REQUIRED)

				find_package(Snappy REQUIRED)

				find_package(RapidJSON REQUIRED)

				find_package(xxHash REQUIRED)

				find_package(yaml-cpp REQUIRED)

				find_package(zstd REQUIRED)

				set(scylla_gen_build_dir "${CMAKE_BINARY_DIR}/gen")

				@@ -138,6 +182,7 @@ target_sources(scylla-main

				    keys.cc

				    multishard_mutation_query.cc

				    mutation_query.cc

				    node_ops/task_manager_module.cc

				    partition_slice_builder.cc

				    querier.cc

				    query.cc

				@@ -151,6 +196,7 @@ target_sources(scylla-main

				    serializer.cc

				    sstables_loader.cc

				    table_helper.cc

				    tasks/task_handler.cc

				    tasks/task_manager.cc

				    timeout_config.cc

				    unimplemented.cc

				@@ -194,6 +240,12 @@ include(check_headers)

				check_headers(check-headers scylla-main

				  GLOB ${CMAKE_CURRENT_SOURCE_DIR}/*.hh)

				option(Scylla_DIST

				  "Build dist targets"

				  ON)

				add_custom_target(compiler-training)

				add_subdirectory(api)

				add_subdirectory(alternator)

				add_subdirectory(db)

				@@ -270,12 +322,20 @@ target_link_libraries(scylla PRIVATE

				    utils)

				target_link_libraries(scylla PRIVATE

				    seastar

				    Seastar::seastar

				    absl::headers

				    yaml-cpp::yaml-cpp

				    Boost::program_options)

				target_include_directories(scylla PRIVATE

				    "${CMAKE_CURRENT_SOURCE_DIR}"

				    "${scylla_gen_build_dir}")

				add_subdirectory(dist)

				add_custom_target(maybe-scylla

				  DEPENDS $<$<CONFIG:Dev>:$<TARGET_FILE:scylla>>)

				add_dependencies(compiler-training

				  maybe-scylla)

				if(Scylla_DIST)

				  add_subdirectory(dist)

				endif()

									
										19

HACKING.md
									
												View File
												
				@@ -19,18 +19,18 @@ $ git submodule update --init --recursive

				### Dependencies

				Scylla is fairly fussy about its build environment, requiring a very recent

				version of the C++20 compiler and numerous tools and libraries to build.

				version of the C++23 compiler and numerous tools and libraries to build.

				Run `./install-dependencies.sh` (as root) to use your Linux distributions's

				package manager to install the appropriate packages on your build machine.

				However, this will only work on very recent distributions. For example,

				currently Fedora users must upgrade to Fedora 32 otherwise the C++ compiler

				will be too old, and not support the new C++20 standard that Scylla uses.

				will be too old, and not support the new C++23 standard that Scylla uses.

				Alternatively, to avoid having to upgrade your build machine or install

				various packages on it, we provide another option - the **frozen toolchain**.

				This is a script, `./tools/toolchain/dbuild`, that can execute build or run

				commands inside a Docker image that contains exactly the right build tools and

				commands inside a container that contains exactly the right build tools and

				libraries. The `dbuild` technique is useful for beginners, but is also the way

				in which ScyllaDB produces official releases, so it is highly recommended.

				@@ -43,6 +43,12 @@ $ ./tools/toolchain/dbuild ninja build/release/scylla

				$ ./tools/toolchain/dbuild ./build/release/scylla --developer-mode 1

				```

				Note: do not mix environemtns - either perform all your work with dbuild, or natively on the host.

				Note2: you can get to an interactive shell within dbuild by running it without any parameters:

				```bash

				$ ./tools/toolchain/dbuild

				```

				### Build system

				**Note**: Compiling Scylla requires, conservatively, 2 GB of memory per native

				@@ -116,6 +122,13 @@ Run all tests through the test execution wrapper with

				$ ./test.py --mode={debug,release}

				```

				or, if you are using `dbuild`, you need to build the code and the tests and then you can run them at will:

				```bash

				$ ./tools/toolchain/dbuild ninja {debug,release,dev}-build

				$ ./tools/toolchain/dbuild ./test.py --mode {debug,release,dev}

				```

				The `--name` argument can be specified to run a particular test.

				Alternatively, you can execute the test executable directly. For example,

									
										10

README.md
									
												View File
												
				@@ -15,7 +15,7 @@ For more information, please see the [ScyllaDB web site].

				## Build Prerequisites

				Scylla is fairly fussy about its build environment, requiring very recent

				versions of the C++20 compiler and of many libraries to build. The document

				versions of the C++23 compiler and of many libraries to build. The document

				[HACKING.md](HACKING.md) includes detailed information on building and

				developing Scylla, but to get Scylla building quickly on (almost) any build

				machine, Scylla offers a [frozen toolchain](tools/toolchain/README.md),

				@@ -84,11 +84,11 @@ Documentation can be found [here](docs/dev/README.md).

				Seastar documentation can be found [here](http://docs.seastar.io/master/index.html).

				User documentation can be found [here](https://docs.scylladb.com/).

				## Training 

				## Training

				Training material and online courses can be found at [Scylla University](https://university.scylladb.com/). 

				The courses are free, self-paced and include hands-on examples. They cover a variety of topics including Scylla data modeling, 

				administration, architecture, basic NoSQL concepts, using drivers for application development, Scylla setup, failover, compactions, 

				Training material and online courses can be found at [Scylla University](https://university.scylladb.com/).

				The courses are free, self-paced and include hands-on examples. They cover a variety of topics including Scylla data modeling,

				administration, architecture, basic NoSQL concepts, using drivers for application development, Scylla setup, failover, compactions,

				multi-datacenters and how Scylla integrates with third-party applications.

				## Contributing to Scylla

4

SCYLLA-VERSION-GEN

View File

@@ -78,7 +78,7 @@ fi
 # Default scylla product/version tags
 PRODUCT=scylla
 VERSION=6.1.0-dev
 VERSION=6.3.0-dev
 if test -f version
 then
@@ -104,7 +104,7 @@ else
 fi
 if [ -f "$OUTPUT_DIR/SCYLLA-RELEASE-FILE" ]; then
 	GIT_COMMIT_FILE=$(cat "$OUTPUT_DIR/SCYLLA-RELEASE-FILE" |cut -d . -f 3)
 	GIT_COMMIT_FILE=$(cat "$OUTPUT_DIR/SCYLLA-RELEASE-FILE" | rev | cut -d . -f 1 | rev)
 	if [ "$GIT_COMMIT" = "$GIT_COMMIT_FILE" ]; then
 		exit 0
 	fi

									
										25

alternator/auth.cc
									
												View File
												
				@@ -8,7 +8,7 @@

				#include "alternator/error.hh"

				#include "auth/common.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include <string>

				#include <string_view>

				#include "bytes.hh"

				@@ -19,6 +19,7 @@

				#include "alternator/executor.hh"

				#include "cql3/selection/selection.hh"

				#include "cql3/result_set.hh"

				#include "types/types.hh"

				#include <seastar/core/coroutine.hh>

				namespace alternator {

				@@ -31,11 +32,12 @@ future<std::string> get_key_from_roles(service::storage_proxy& proxy, auth::serv

				    dht::partition_range_vector partition_ranges{dht::partition_range(dht::decorate_key(*schema, pk))};

				    std::vector<query::clustering_range> bounds{query::clustering_range::make_open_ended_both_sides()};

				    const column_definition* salted_hash_col = schema->get_column_definition(bytes("salted_hash"));

				    if (!salted_hash_col) {

				        co_await coroutine::return_exception(api_error::unrecognized_client(format("Credentials cannot be fetched for: {}", username)));

				    const column_definition* can_login_col = schema->get_column_definition(bytes("can_login"));

				    if (!salted_hash_col || !can_login_col) {

				        co_await coroutine::return_exception(api_error::unrecognized_client(fmt::format("Credentials cannot be fetched for: {}", username)));

				    }

				    auto selection = cql3::selection::selection::for_columns(schema, {salted_hash_col});

				    auto partition_slice = query::partition_slice(std::move(bounds), {}, query::column_id_vector{salted_hash_col->id}, selection->get_query_options());

				    auto selection = cql3::selection::selection::for_columns(schema, {salted_hash_col, can_login_col});

				    auto partition_slice = query::partition_slice(std::move(bounds), {}, query::column_id_vector{salted_hash_col->id, can_login_col->id}, selection->get_query_options());

				    auto command = ::make_lw_shared<query::read_command>(schema->id(), schema->version(), partition_slice,

				            proxy.get_max_result_size(partition_slice), query::tombstone_limit(proxy.get_tombstone_limit()));

				    auto cl = auth::password_authenticator::consistency_for_user(username);

				@@ -49,11 +51,18 @@ future<std::string> get_key_from_roles(service::storage_proxy& proxy, auth::serv

				    auto result_set = builder.build();

				    if (result_set->empty()) {

				        co_await coroutine::return_exception(api_error::unrecognized_client(format("User not found: {}", username)));

				        co_await coroutine::return_exception(api_error::unrecognized_client(fmt::format("User not found: {}", username)));

				    }

				    const managed_bytes_opt& salted_hash = result_set->rows().front().front(); // We only asked for 1 row and 1 column

				    const auto& result = result_set->rows().front();

				    bool can_login = result[1] && value_cast<bool>(boolean_type->deserialize(*result[1]));

				    if (!can_login) {

				        // This is a valid role name, but has "login=False" so should not be

				        // usable for authentication (see #19735).

				        co_await coroutine::return_exception(api_error::unrecognized_client(fmt::format("Role {} has login=false so cannot be used for login", username)));

				    }

				    const managed_bytes_opt& salted_hash = result.front();

				    if (!salted_hash) {

				        co_await coroutine::return_exception(api_error::unrecognized_client(format("No password found for user: {}", username)));

				        co_await coroutine::return_exception(api_error::unrecognized_client(fmt::format("No password found for user: {}", username)));

				    }

				    co_return value_cast<sstring>(utf8_type->deserialize(*salted_hash));

				}

									
										14

alternator/conditions.cc
									
												View File
												
				@@ -15,8 +15,6 @@

				#include "utils/base64.hh"

				#include "utils/rjson.hh"

				#include <stdexcept>

				#include <boost/algorithm/cxx11/all_of.hpp>

				#include <boost/algorithm/cxx11/any_of.hpp>

				#include "utils/overloaded_functor.hh"

				#include "expressions.hh"

				@@ -42,12 +40,12 @@ comparison_operator_type get_comparison_operator(const rjson::value& comparison_

				            {"NOT_CONTAINS", comparison_operator_type::NOT_CONTAINS},

				    };

				    if (!comparison_operator.IsString()) {

				        throw api_error::validation(format("Invalid comparison operator definition {}", rjson::print(comparison_operator)));

				        throw api_error::validation(fmt::format("Invalid comparison operator definition {}", rjson::print(comparison_operator)));

				    }

				    std::string op = comparison_operator.GetString();

				    auto it = ops.find(op);

				    if (it == ops.end()) {

				        throw api_error::validation(format("Unsupported comparison operator {}", op));

				        throw api_error::validation(fmt::format("Unsupported comparison operator {}", op));

				    }

				    return it->second;

				}

				@@ -429,7 +427,7 @@ static bool check_BETWEEN(const T& v, const T& lb, const T& ub, bool bounds_from

				    if (cmp_lt()(ub, lb)) {

				        if (bounds_from_query) {

				            throw api_error::validation(

				                format("BETWEEN operator requires lower_bound <= upper_bound, but {} > {}", lb, ub));

				                fmt::format("BETWEEN operator requires lower_bound <= upper_bound, but {} > {}", lb, ub));

				        } else {

				            return false;

				        }

				@@ -613,7 +611,7 @@ conditional_operator_type get_conditional_operator(const rjson::value& req) {

				        return conditional_operator_type::OR;

				    } else {

				        throw api_error::validation(

				                format("'ConditionalOperator' parameter must be AND, OR or missing. Found {}.", s));

				                fmt::format("'ConditionalOperator' parameter must be AND, OR or missing. Found {}.", s));

				    }

				}

				@@ -743,9 +741,9 @@ bool verify_condition_expression(

				            };

				            switch (list.op) {

				            case '&':

				                return boost::algorithm::all_of(list.conditions, verify_condition);

				                return std::ranges::all_of(list.conditions, verify_condition);

				            case '|':

				                return boost::algorithm::any_of(list.conditions, verify_condition);

				                return std::ranges::any_of(list.conditions, verify_condition);

				            default:

				                // Shouldn't happen unless we have a bug in the parser

				                throw std::logic_error("bad operator in condition_list");

									
										6

alternator/controller.cc
									
												View File
												
				@@ -130,10 +130,10 @@ future<> controller::start_server() {

				                std::throw_with_nested(std::runtime_error("Failed to set up Alternator TLS credentials"));

				            }

				        }

				        bool alternator_enforce_authorization = _config.alternator_enforce_authorization();

				        _server.invoke_on_all(

				                [this, addr, alternator_port, alternator_https_port, creds = std::move(creds), alternator_enforce_authorization] (server& server) mutable {

				            return server.init(addr, alternator_port, alternator_https_port, creds, alternator_enforce_authorization,

				                [this, addr, alternator_port, alternator_https_port, creds = std::move(creds)] (server& server) mutable {

				            return server.init(addr, alternator_port, alternator_https_port, creds,

				                    _config.alternator_enforce_authorization,

				                    &_memory_limiter.local().get_semaphore(),

				                    _config.max_concurrent_requests_per_shard);

				        }).handle_exception([this, addr, alternator_port, alternator_https_port] (std::exception_ptr ep) {

502

alternator/executor.cc

View File

File diff suppressed because it is too large Load Diff

									
										15

alternator/executor.hh
									
												View File
												
				@@ -23,6 +23,8 @@

				#include "utils/rjson.hh"

				#include "utils/updateable_value.hh"

				#include "tracing/trace_state.hh"

				namespace db {

				    class system_distributed_keyspace;

				}

				@@ -50,6 +52,8 @@ class gossiper;

				}

				class schema_builder;

				namespace alternator {

				class rmw_operation;

				@@ -158,6 +162,7 @@ class executor : public peering_sharded_service<executor> {

				    service::migration_manager& _mm;

				    db::system_distributed_keyspace& _sdks;

				    cdc::metadata& _cdc_metadata;

				    utils::updateable_value<bool> _enforce_authorization;

				    // An smp_service_group to be used for limiting the concurrency when

				    // forwarding Alternator request between shards - if necessary for LWT.

				    smp_service_group _ssg;

				@@ -176,10 +181,7 @@ public:

				             db::system_distributed_keyspace& sdks,

				             cdc::metadata& cdc_metadata,

				             smp_service_group ssg,

				             utils::updateable_value<uint32_t> default_timeout_in_ms)

				        : _gossiper(gossiper), _proxy(proxy), _mm(mm), _sdks(sdks), _cdc_metadata(cdc_metadata), _ssg(ssg) {

				        s_default_timeout_in_ms = std::move(default_timeout_in_ms);

				    }

				             utils::updateable_value<uint32_t> default_timeout_in_ms);

				    future<request_return_type> create_table(client_state& client_state, tracing::trace_state_ptr trace_state, service_permit permit, rjson::value request);

				    future<request_return_type> describe_table(client_state& client_state, tracing::trace_state_ptr trace_state, service_permit permit, rjson::value request);

				@@ -262,4 +264,9 @@ public:

				// add more than a couple of levels in its own output construction.

				bool is_big(const rjson::value& val, int big_size = 100'000);

				// Check CQL's Role-Based Access Control (RBAC) permission (MODIFY,

				// SELECT, DROP, etc.) on the given table. When permission is denied an

				// appropriate user-readable api_error::access_denied is thrown.

				future<> verify_permission(bool enforce_authorization, const service::client_state&, const schema_ptr&, auth::permission);

				}

									
										19

alternator/expressions.cc
									
												View File
												
				@@ -20,9 +20,6 @@

				#include <seastar/core/print.hh>

				#include <seastar/util/log.hh>

				#include <boost/algorithm/cxx11/any_of.hpp>

				#include <boost/algorithm/cxx11/all_of.hpp>

				#include <functional>

				#include <unordered_map>

				@@ -57,10 +54,10 @@ static Result parse(const char* input_name, std::string_view input, Func&& f) {

				        // TODO: displayRecognitionError could set a position inside the

				        // expressions_syntax_error in throws, and we could use it here to

				        // mark the broken position in 'input'.

				        throw expressions_syntax_error(format("Failed parsing {} '{}': {}",

				        throw expressions_syntax_error(fmt::format("Failed parsing {} '{}': {}",

				            input_name, input, e.what()));

				    } catch (...) {

				        throw expressions_syntax_error(format("Failed parsing {} '{}': {}",

				        throw expressions_syntax_error(fmt::format("Failed parsing {} '{}': {}",

				            input_name, input, std::current_exception()));

				    }

				}

				@@ -160,12 +157,12 @@ static std::optional<std::string> resolve_path_component(const std::string& colu

				    if (column_name.size() > 0 && column_name.front() == '#') {

				        if (!expression_attribute_names) {

				            throw api_error::validation(

				                    format("ExpressionAttributeNames missing, entry '{}' required by expression", column_name));

				                    fmt::format("ExpressionAttributeNames missing, entry '{}' required by expression", column_name));

				        }

				        const rjson::value* value = rjson::find(*expression_attribute_names, column_name);

				        if (!value || !value->IsString()) {

				            throw api_error::validation(

				                    format("ExpressionAttributeNames missing entry '{}' required by expression", column_name));

				                    fmt::format("ExpressionAttributeNames missing entry '{}' required by expression", column_name));

				        }

				        used_attribute_names.emplace(column_name);

				        return std::string(rjson::to_string_view(*value));

				@@ -202,16 +199,16 @@ static void resolve_constant(parsed::constant& c,

				        [&] (const std::string& valref) {

				            if (!expression_attribute_values) {

				                throw api_error::validation(

				                        format("ExpressionAttributeValues missing, entry '{}' required by expression", valref));

				                        fmt::format("ExpressionAttributeValues missing, entry '{}' required by expression", valref));

				            }

				            const rjson::value* value = rjson::find(*expression_attribute_values, valref);

				            if (!value) {

				                throw api_error::validation(

				                        format("ExpressionAttributeValues missing entry '{}' required by expression", valref));

				                        fmt::format("ExpressionAttributeValues missing entry '{}' required by expression", valref));

				            }

				            if (value->IsNull()) {

				                throw api_error::validation(

				                        format("ExpressionAttributeValues null value for entry '{}' required by expression", valref));

				                        fmt::format("ExpressionAttributeValues null value for entry '{}' required by expression", valref));

				            }

				            validate_value(*value, "ExpressionAttributeValues");

				            used_attribute_values.emplace(valref);

				@@ -708,7 +705,7 @@ rjson::value calculate_value(const parsed::value& v,

				            auto function_it = function_handlers.find(std::string_view(f._function_name));

				            if (function_it == function_handlers.end()) {

				                throw api_error::validation(

				                        format("{}: unknown function '{}' called.", caller, f._function_name));

				                        fmt::format("{}: unknown function '{}' called.", caller, f._function_name));

				            }

				            return function_it->second(caller, previous_item, f);

				        },

									
										2

alternator/rmw_operation.hh
									
												View File
												
				@@ -12,6 +12,8 @@

				#include "service/paxos/cas_request.hh"

				#include "utils/rjson.hh"

				#include "executor.hh"

				#include "tracing/trace_state.hh"

				#include "keys.hh"

				namespace alternator {

									
										37

alternator/serialization.cc
									
												View File
												
				@@ -8,7 +8,7 @@

				#include "utils/base64.hh"

				#include "utils/rjson.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "serialization.hh"

				#include "error.hh"

				#include "concrete_types.hh"

				@@ -143,17 +143,17 @@ static big_decimal parse_and_validate_number(std::string_view s) {

				        big_decimal ret(s);

				        auto [magnitude, precision] = internal::get_magnitude_and_precision(s);

				        if (magnitude > 125) {

				            throw api_error::validation(format("Number overflow: {}. Attempting to store a number with magnitude larger than supported range.", s));

				            throw api_error::validation(fmt::format("Number overflow: {}. Attempting to store a number with magnitude larger than supported range.", s));

				        }

				        if (magnitude < -130) {

				            throw api_error::validation(format("Number underflow: {}. Attempting to store a number with magnitude lower than supported range.", s));

				            throw api_error::validation(fmt::format("Number underflow: {}. Attempting to store a number with magnitude lower than supported range.", s));

				        }

				        if (precision > 38) {

				            throw api_error::validation(format("Number too precise: {}. Attempting to store a number with more significant digits than supported.", s));

				            throw api_error::validation(fmt::format("Number too precise: {}. Attempting to store a number with more significant digits than supported.", s));

				        }

				        return ret;

				    } catch (const marshal_exception& e) {

				        throw api_error::validation(format("The parameter cannot be converted to a numeric value: {}", s));

				        throw api_error::validation(fmt::format("The parameter cannot be converted to a numeric value: {}", s));

				    }

				}

				@@ -265,7 +265,7 @@ bytes get_key_column_value(const rjson::value& item, const column_definition& co

				    std::string column_name = column.name_as_text();

				    const rjson::value* key_typed_value = rjson::find(item, column_name);

				    if (!key_typed_value) {

				        throw api_error::validation(format("Key column {} not found", column_name));

				        throw api_error::validation(fmt::format("Key column {} not found", column_name));

				    }

				    return get_key_from_typed_value(*key_typed_value, column);

				}

				@@ -277,19 +277,26 @@ bytes get_key_column_value(const rjson::value& item, const column_definition& co

				// mentioned in the exception message).

				// If the type does match, a reference to the encoded value is returned.

				static const rjson::value& get_typed_value(const rjson::value& key_typed_value, std::string_view type_str, std::string_view name, std::string_view value_name) {

				    if (!key_typed_value.IsObject() || key_typed_value.MemberCount() != 1 ||

				            !key_typed_value.MemberBegin()->value.IsString()) {

				    if (!key_typed_value.IsObject() || key_typed_value.MemberCount() != 1) {

				        throw api_error::validation(

				                format("Malformed value object for {} {}: {}",

				                fmt::format("Malformed value object for {} {}: {}",

				                        value_name, name, key_typed_value));

				    }

				    auto it = key_typed_value.MemberBegin();

				    if (rjson::to_string_view(it->name) != type_str) {

				        throw api_error::validation(

				                format("Type mismatch: expected type {} for {} {}, got type {}",

				                fmt::format("Type mismatch: expected type {} for {} {}, got type {}",

				                        type_str, value_name, name, it->name));

				    }

				    // We assume this function is called just for key types (S, B, N), and

				    // all of those always have a string value in the JSON.

				    if (!it->value.IsString()) {

				        throw api_error::validation(

				            fmt::format("Malformed value object for {} {}: {}",

				                    value_name, name, key_typed_value));

				    }

				    return it->value;

				}

				@@ -395,16 +402,16 @@ position_in_partition pos_from_json(const rjson::value& item, schema_ptr schema)

				big_decimal unwrap_number(const rjson::value& v, std::string_view diagnostic) {

				    if (!v.IsObject() || v.MemberCount() != 1) {

				        throw api_error::validation(format("{}: invalid number object", diagnostic));

				        throw api_error::validation(fmt::format("{}: invalid number object", diagnostic));

				    }

				    auto it = v.MemberBegin();

				    if (it->name != "N") {

				        throw api_error::validation(format("{}: expected number, found type '{}'", diagnostic, it->name));

				        throw api_error::validation(fmt::format("{}: expected number, found type '{}'", diagnostic, it->name));

				    }

				    if (!it->value.IsString()) {

				        // We shouldn't reach here. Callers normally validate their input

				        // earlier with validate_value().

				        throw api_error::validation(format("{}: improperly formatted number constant", diagnostic));

				        throw api_error::validation(fmt::format("{}: improperly formatted number constant", diagnostic));

				    }

				    big_decimal ret = parse_and_validate_number(rjson::to_string_view(it->value));

				    return ret;

				@@ -485,7 +492,7 @@ rjson::value set_sum(const rjson::value& v1, const rjson::value& v2) {

				    auto [set1_type, set1] = unwrap_set(v1);

				    auto [set2_type, set2] = unwrap_set(v2);

				    if (set1_type != set2_type) {

				        throw api_error::validation(format("Mismatched set types: {} and {}", set1_type, set2_type));

				        throw api_error::validation(fmt::format("Mismatched set types: {} and {}", set1_type, set2_type));

				    }

				    if (!set1 || !set2) {

				        throw api_error::validation("UpdateExpression: ADD operation for sets must be given sets as arguments");

				@@ -513,7 +520,7 @@ std::optional<rjson::value> set_diff(const rjson::value& v1, const rjson::value&

				    auto [set1_type, set1] = unwrap_set(v1);

				    auto [set2_type, set2] = unwrap_set(v2);

				    if (set1_type != set2_type) {

				        throw api_error::validation(format("Set DELETE type mismatch: {} and {}", set1_type, set2_type));

				        throw api_error::validation(fmt::format("Set DELETE type mismatch: {} and {}", set1_type, set2_type));

				    }

				    if (!set1 || !set2) {

				        throw api_error::validation("UpdateExpression: DELETE operation can only be performed on a set");

									
										61

alternator/server.cc
									
												View File
												
				@@ -7,7 +7,8 @@

				 */

				#include "alternator/server.hh"

				#include "log.hh"

				#include "gms/application_state.hh"

				#include "utils/log.hh"

				#include <fmt/ranges.h>

				#include <seastar/http/function_handlers.hh>

				#include <seastar/http/short_streams.hh>

				@@ -17,7 +18,10 @@

				#include <seastar/util/short_streams.hh>

				#include "seastarx.hh"

				#include "error.hh"

				#include "service/client_state.hh"

				#include "service/qos/service_level_controller.hh"

				#include "utils/assert.hh"

				#include "timeout_config.hh"

				#include "utils/rjson.hh"

				#include "auth.hh"

				#include <cctype>

				@@ -208,10 +212,32 @@ protected:

				        // using _gossiper().get_live_members(). But getting

				        // just the list of live nodes in this DC needs more elaborate code:

				        auto& topology = _proxy.get_token_metadata_ptr()->get_topology();

				        sstring local_dc = topology.get_datacenter();

				        std::unordered_set<gms::inet_address> local_dc_nodes = topology.get_datacenter_endpoints().at(local_dc);

				        // /localnodes lists nodes in a single DC. By default the DC of this

				        // server is used, but it can be overridden by a "dc" query option.

				        // If the DC does not exist, we return an empty list - not an error.

				        sstring query_dc = req->get_query_param("dc");

				        sstring local_dc = query_dc.empty() ? topology.get_datacenter() : query_dc;

				        std::unordered_set<gms::inet_address> local_dc_nodes;

				        const auto& endpoints = topology.get_datacenter_endpoints();

				        auto dc_it = endpoints.find(local_dc);

				        if (dc_it != endpoints.end()) {

				            local_dc_nodes = dc_it->second;

				        }

				        // By default, /localnodes lists the nodes of all racks in the given

				        // DC, unless a single rack is selected by the "rack" query option.

				        // If the rack does not exist, we return an empty list - not an error.

				        sstring query_rack = req->get_query_param("rack");

				        for (auto& ip : local_dc_nodes) {

				            if (_gossiper.is_alive(ip)) {

				            if (!query_rack.empty()) {

				                auto rack = _gossiper.get_application_state_value(ip, gms::application_state::RACK);

				                if (rack != query_rack) {

				                    continue;

				                }

				            }

				            // Note that it's not enough for the node to be is_alive() - a

				            // node joining the cluster is also "alive" but not responsive to

				            // requests. We need the node to be in normal state. See #19694.

				            if (_gossiper.is_normal(ip)) {

				                // Use the gossiped broadcast_rpc_address if available instead

				                // of the internal IP address "ip". See discussion in #18711.

				                rjson::push_back(results, rjson::from_string(_gossiper.get_rpc_address(ip)));

				@@ -257,7 +283,7 @@ future<std::string> server::verify_signature(const request& req, const chunked_c

				    std::string_view authorization_header = authorization_it->second;

				    auto pos = authorization_header.find_first_of(' ');

				    if (pos == std::string_view::npos || authorization_header.substr(0, pos) != "AWS4-HMAC-SHA256") {

				        throw api_error::invalid_signature(format("Authorization header must use AWS4-HMAC-SHA256 algorithm: {}", authorization_header));

				        throw api_error::invalid_signature(fmt::format("Authorization header must use AWS4-HMAC-SHA256 algorithm: {}", authorization_header));

				    }

				    authorization_header.remove_prefix(pos+1);

				    std::string credential;

				@@ -292,7 +318,7 @@ future<std::string> server::verify_signature(const request& req, const chunked_c

				    std::vector<std::string_view> credential_split = split(credential, '/');

				    if (credential_split.size() != 5) {

				        throw api_error::validation(format("Incorrect credential information format: {}", credential));

				        throw api_error::validation(fmt::format("Incorrect credential information format: {}", credential));

				    }

				    std::string user(credential_split[0]);

				    std::string datestamp(credential_split[1]);

				@@ -377,7 +403,7 @@ static tracing::trace_state_ptr maybe_trace_query(service::client_state& client_

				        std::string buf;

				        tracing::add_session_param(trace_state, "alternator_op", op);

				        tracing::add_query(trace_state, truncated_content_view(query, buf));

				        tracing::begin(trace_state, format("Alternator {}", op), client_state.get_client_address());

				        tracing::begin(trace_state, seastar::format("Alternator {}", op), client_state.get_client_address());

				        if (!username.empty()) {

				            tracing::set_username(trace_state, auth::authenticated_user(username));

				        }

				@@ -402,7 +428,7 @@ future<executor::request_return_type> server::handle_api_request(std::unique_ptr

				        ++_executor._stats.requests_blocked_memory;

				    }

				    auto units = co_await std::move(units_fut);

				    assert(req->content_stream);

				    SCYLLA_ASSERT(req->content_stream);

				    chunked_content content = co_await util::read_entire_stream(*req->content_stream);

				    auto username = co_await verify_signature(*req, content);

				@@ -413,7 +439,7 @@ future<executor::request_return_type> server::handle_api_request(std::unique_ptr

				    auto callback_it = _callbacks.find(op);

				    if (callback_it == _callbacks.end()) {

				        _executor._stats.unsupported_operations++;

				        co_return api_error::unknown_operation(format("Unsupported operation {}", op));

				        co_return api_error::unknown_operation(fmt::format("Unsupported operation {}", op));

				    }

				    if (_pending_requests.get_count() >= _max_concurrent_requests) {

				        _executor._stats.requests_shed++;

				@@ -421,11 +447,11 @@ future<executor::request_return_type> server::handle_api_request(std::unique_ptr

				    }

				    _pending_requests.enter();

				    auto leave = defer([this] () noexcept { _pending_requests.leave(); });

				    //FIXME: Client state can provide more context, e.g. client's endpoint address

				    // We use unique_ptr because client_state cannot be moved or copied

				    executor::client_state client_state = username.empty()

				        ? service::client_state{service::client_state::internal_tag()}

				        : service::client_state{service::client_state::internal_tag(), _auth_service, _sl_controller, username};

				    executor::client_state client_state(service::client_state::external_tag(),

				        _auth_service, &_sl_controller, _timeout_config.current_values(), req->get_client_address());

				    if (!username.empty()) {

				        client_state.set_login(auth::authenticated_user(username));

				    }

				    co_await client_state.maybe_update_per_service_level_params();

				    tracing::trace_state_ptr trace_state = maybe_trace_query(client_state, username, op, content);

				@@ -472,6 +498,7 @@ server::server(executor& exec, service::storage_proxy& proxy, gms::gossiper& gos

				        , _enforce_authorization(false)

				        , _enabled_servers{}

				        , _pending_requests{}

				        , _timeout_config(_proxy.data_dictionary().get_config())

				      , _callbacks{

				        {"CreateTable", [] (executor& e, executor::client_state& client_state, tracing::trace_state_ptr trace_state, service_permit permit, rjson::value json_request, std::unique_ptr<request> req) {

				            return e.create_table(client_state, std::move(trace_state), std::move(permit), std::move(json_request));

				@@ -549,9 +576,9 @@ server::server(executor& exec, service::storage_proxy& proxy, gms::gossiper& gos

				}

				future<> server::init(net::inet_address addr, std::optional<uint16_t> port, std::optional<uint16_t> https_port, std::optional<tls::credentials_builder> creds,

				        bool enforce_authorization, semaphore* memory_limiter, utils::updateable_value<uint32_t> max_concurrent_requests) {

				        utils::updateable_value<bool> enforce_authorization, semaphore* memory_limiter, utils::updateable_value<uint32_t> max_concurrent_requests) {

				    _memory_limiter = memory_limiter;

				    _enforce_authorization = enforce_authorization;

				    _enforce_authorization = std::move(enforce_authorization);

				    _max_concurrent_requests = std::move(max_concurrent_requests);

				    if (!port && !https_port) {

				        return make_exception_future<>(std::runtime_error("Either regular port or TLS port"

				@@ -636,7 +663,7 @@ future<> server::json_parser::stop() {

				const char* api_error::what() const noexcept {

				    if (_what_string.empty()) {

				        _what_string = format("{} {}: {}", std::to_underlying(_http_code), _type, _msg);

				        _what_string = fmt::format("{} {}: {}", std::to_underlying(_http_code), _type, _msg);

				    }

				    return _what_string.c_str();

				}

									
										9

alternator/server.hh
									
												View File
												
				@@ -39,9 +39,14 @@ class server {

				    qos::service_level_controller& _sl_controller;

				    key_cache _key_cache;

				    bool _enforce_authorization;

				    utils::updateable_value<bool> _enforce_authorization;

				    utils::small_vector<std::reference_wrapper<seastar::httpd::http_server>, 2> _enabled_servers;

				    gate _pending_requests;

				    // In some places we will need a CQL updateable_timeout_config object even

				    // though it isn't really relevant for Alternator which defines its own

				    // timeouts separately. We can create this object only once.

				    updateable_timeout_config _timeout_config;

				    alternator_callbacks_map _callbacks;

				    semaphore* _memory_limiter;

				@@ -71,7 +76,7 @@ public:

				    server(executor& executor, service::storage_proxy& proxy, gms::gossiper& gossiper, auth::service& service, qos::service_level_controller& sl_controller);

				    future<> init(net::inet_address addr, std::optional<uint16_t> port, std::optional<uint16_t> https_port, std::optional<tls::credentials_builder> creds,

				            bool enforce_authorization, semaphore* memory_limiter, utils::updateable_value<uint32_t> max_concurrent_requests);

				            utils::updateable_value<bool> enforce_authorization, semaphore* memory_limiter, utils::updateable_value<uint32_t> max_concurrent_requests);

				    future<> stop();

				private:

				    void set_routes(seastar::httpd::routes& r);

									
										6

alternator/stats.cc
									
												View File
												
				@@ -67,6 +67,8 @@ stats::stats() : api_operations{} {

				            OPERATION_LATENCY(get_item_latency, "GetItem")

				            OPERATION_LATENCY(delete_item_latency, "DeleteItem")

				            OPERATION_LATENCY(update_item_latency, "UpdateItem")

				            OPERATION_LATENCY(batch_write_item_latency, "BatchWriteItem")

				            OPERATION_LATENCY(batch_get_item_latency, "BatchGetItem")

				            OPERATION(list_streams, "ListStreams")

				            OPERATION(describe_stream, "DescribeStream")

				            OPERATION(get_shard_iterator, "GetShardIterator")

				@@ -94,6 +96,10 @@ stats::stats() : api_operations{} {

				                    seastar::metrics::description("number of rows read and matched during filtering operations")),

				            seastar::metrics::make_total_operations("filtered_rows_dropped_total", [this] { return cql_stats.filtered_rows_read_total - cql_stats.filtered_rows_matched_total; },

				                    seastar::metrics::description("number of rows read and dropped during filtering operations")),

				                    seastar::metrics::make_counter("batch_item_count", seastar::metrics::description("The total number of items processed across all batches"),{op("BatchWriteItem")},

				                            api_operations.batch_write_item_batch_total).set_skip_when_empty(),

				                    seastar::metrics::make_counter("batch_item_count", seastar::metrics::description("The total number of items processed across all batches"),{op("BatchGetItem")},

				                            api_operations.batch_get_item_batch_total).set_skip_when_empty(),

				    });

				}

									
										4

alternator/stats.hh
									
												View File
												
				@@ -26,6 +26,8 @@ public:

				    struct {

				        uint64_t batch_get_item = 0;

				        uint64_t batch_write_item = 0;

				        uint64_t batch_get_item_batch_total = 0;

				        uint64_t batch_write_item_batch_total = 0;

				        uint64_t create_backup = 0;

				        uint64_t create_global_table = 0;

				        uint64_t create_table = 0;

				@@ -69,6 +71,8 @@ public:

				        utils::timed_rate_moving_average_summary_and_histogram get_item_latency;

				        utils::timed_rate_moving_average_summary_and_histogram delete_item_latency;

				        utils::timed_rate_moving_average_summary_and_histogram update_item_latency;

				        utils::timed_rate_moving_average_summary_and_histogram batch_write_item_latency;

				        utils::timed_rate_moving_average_summary_and_histogram batch_get_item_latency;

				        utils::timed_rate_moving_average_summary_and_histogram get_records_latency;

				    } api_operations;

				    // Miscellaneous event counters

									
										32

alternator/streams.cc
									
												View File
												
				@@ -13,6 +13,7 @@

				#include <seastar/json/formatter.hh>

				#include "auth/permission.hh"

				#include "db/config.hh"

				#include "cdc/log.hh"

				@@ -818,11 +819,13 @@ future<executor::request_return_type> executor::get_records(client_state& client

				    }

				    if (!schema || !base || !is_alternator_keyspace(schema->ks_name())) {

				        throw api_error::resource_not_found(fmt::to_string(iter.table));

				        co_return api_error::resource_not_found(fmt::to_string(iter.table));

				    }

				    tracing::add_table_name(trace_state, schema->ks_name(), schema->cf_name());

				    co_await verify_permission(_enforce_authorization, client_state, schema, auth::permission::SELECT);

				    db::consistency_level cl = db::consistency_level::LOCAL_QUORUM;

				    partition_key pk = iter.shard.id.to_partition_key(*schema);

				@@ -841,19 +844,21 @@ future<executor::request_return_type> executor::get_records(client_state& client

				    static const bytes op_column_name = cdc::log_meta_column_name_bytes("operation");

				    static const bytes eor_column_name = cdc::log_meta_column_name_bytes("end_of_batch");

				    std::optional<attrs_to_get> key_names = boost::copy_range<attrs_to_get>(

				        boost::range::join(std::move(base->partition_key_columns()), std::move(base->clustering_key_columns()))

				        | boost::adaptors::transformed([&] (const column_definition& cdef) {

				    std::optional<attrs_to_get> key_names =

				        base->primary_key_columns()

				        | std::views::transform([&] (const column_definition& cdef) {

				            return std::make_pair<std::string, attrs_to_get_node>(cdef.name_as_text(), {}); })

				    );

				        | std::ranges::to<attrs_to_get>()

				    ;

				    // Include all base table columns as values (in case pre or post is enabled).

				    // This will include attributes not stored in the frozen map column

				    std::optional<attrs_to_get> attr_names = boost::copy_range<attrs_to_get>(base->regular_columns()

				    std::optional<attrs_to_get> attr_names = base->regular_columns()

				        // this will include the :attrs column, which we will also force evaluating. 

				        // But not having this set empty forces out any cdc columns from actual result 

				        | boost::adaptors::transformed([] (const column_definition& cdef) {

				        | std::views::transform([] (const column_definition& cdef) {

				            return std::make_pair<std::string, attrs_to_get_node>(cdef.name_as_text(), {}); })

				    );

				        | std::ranges::to<attrs_to_get>()

				    ;

				    std::vector<const column_definition*> columns;

				    columns.reserve(schema->all_columns().size());

				@@ -864,10 +869,11 @@ future<executor::request_return_type> executor::get_records(client_state& client

				    std::transform(pks.begin(), pks.end(), std::back_inserter(columns), [](auto& c) { return &c; });

				    std::transform(cks.begin(), cks.end(), std::back_inserter(columns), [](auto& c) { return &c; });

				    auto regular_columns = boost::copy_range<query::column_id_vector>(schema->regular_columns() 

				        | boost::adaptors::filtered([](const column_definition& cdef) { return cdef.name() == op_column_name || cdef.name() == eor_column_name || !cdc::is_cdc_metacolumn_name(cdef.name_as_text()); })

				        | boost::adaptors::transformed([&] (const column_definition& cdef) { columns.emplace_back(&cdef); return cdef.id; })

				    );

				    auto regular_columns = schema->regular_columns()

				        | std::views::filter([](const column_definition& cdef) { return cdef.name() == op_column_name || cdef.name() == eor_column_name || !cdc::is_cdc_metacolumn_name(cdef.name_as_text()); })

				        | std::views::transform([&] (const column_definition& cdef) { columns.emplace_back(&cdef); return cdef.id; })

				        | std::ranges::to<query::column_id_vector>()

				    ;

				    stream_view_type type = cdc_options_to_steam_view_type(base->cdc_options());

				@@ -887,7 +893,7 @@ future<executor::request_return_type> executor::get_records(client_state& client

				    auto command = ::make_lw_shared<query::read_command>(schema->id(), schema->version(), partition_slice, _proxy.get_max_result_size(partition_slice),

				            query::tombstone_limit(_proxy.get_tombstone_limit()), query::row_limit(limit * mul));

				    return _proxy.query(schema, std::move(command), std::move(partition_ranges), cl, service::storage_proxy::coordinator_query_options(default_timeout(), std::move(permit), client_state)).then(

				    co_return co_await _proxy.query(schema, std::move(command), std::move(partition_ranges), cl, service::storage_proxy::coordinator_query_options(default_timeout(), std::move(permit), client_state)).then(

				            [this, schema, partition_slice = std::move(partition_slice), selection = std::move(selection), start_time = std::move(start_time), limit, key_names = std::move(key_names), attr_names = std::move(attr_names), type, iter, high_ts] (service::storage_proxy::coordinator_query_result qr) mutable {       

				        cql3::selection::result_set_builder builder(*selection, gc_clock::now());

				        query::result_view::consume(*qr.query_result, partition_slice, cql3::selection::result_set_builder::visitor(builder, *schema, *selection));

									
										114

alternator/ttl.cc
									
												View File
												
				@@ -23,9 +23,10 @@

				#include "gms/inet_address.hh"

				#include "inet_address_vectors.hh"

				#include "locator/abstract_replication_strategy.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "gc_clock.hh"

				#include "replica/database.hh"

				#include "service/client_state.hh"

				#include "service_permit.hh"

				#include "timestamp.hh"

				#include "service/storage_proxy.hh"

				@@ -35,6 +36,7 @@

				#include "mutation/mutation.hh"

				#include "types/types.hh"

				#include "types/map.hh"

				#include "utils/assert.hh"

				#include "utils/rjson.hh"

				#include "utils/big_decimal.hh"

				#include "cql3/selection/selection.hh"

				@@ -97,6 +99,7 @@ future<executor::request_return_type> executor::update_time_to_live(client_state

				    }

				    sstring attribute_name(v->GetString(), v->GetStringLength());

				    co_await verify_permission(_enforce_authorization, client_state, schema, auth::permission::ALTER);

				    co_await db::modify_tags(_mm, schema->ks_name(), schema->cf_name(), [&](std::map<sstring, sstring>& tags_map) {

				        if (enabled) {

				            if (tags_map.contains(TTL_TAG_KEY)) {

				@@ -312,7 +315,7 @@ static size_t random_offset(size_t min, size_t max) {

				// this range's primary node is down. For this we need to return not just

				// a list of this node's secondary ranges - but also the primary owner of

				// each of those ranges.

				static std::vector<std::pair<dht::token_range, gms::inet_address>> get_secondary_ranges(

				static future<std::vector<std::pair<dht::token_range, gms::inet_address>>> get_secondary_ranges(

				        const locator::effective_replication_map_ptr& erm,

				        gms::inet_address ep) {

				    const auto& tm = *erm->get_token_metadata_ptr();

				@@ -323,6 +326,7 @@ static std::vector<std::pair<dht::token_range, gms::inet_address>> get_secondary

				    }

				    auto prev_tok = sorted_tokens.back();

				    for (const auto& tok : sorted_tokens) {

				        co_await coroutine::maybe_yield();

				        inet_address_vector_replica_set eps = erm->get_natural_endpoints(tok);

				        if (eps.size() <= 1 || eps[1] != ep) {

				            prev_tok = tok;

				@@ -350,7 +354,7 @@ static std::vector<std::pair<dht::token_range, gms::inet_address>> get_secondary

				        }

				        prev_tok = tok;

				    }

				    return ret;

				    co_return ret;

				}

				@@ -386,63 +390,63 @@ static std::vector<std::pair<dht::token_range, gms::inet_address>> get_secondary

				//

				// FIXME: Check if this algorithm is safe with tablet migration.

				// https://github.com/scylladb/scylladb/issues/16567

				enum primary_or_secondary_t {primary, secondary};

				template<primary_or_secondary_t primary_or_secondary>

				class token_ranges_owned_by_this_shard {

				    // ranges_holder_primary holds just the primary ranges themselves

				    class ranges_holder_primary {

				        const dht::token_range_vector _token_ranges;

				     public:

				        ranges_holder_primary(const locator::vnode_effective_replication_map_ptr& erm, gms::gossiper& g, gms::inet_address ep)

				            : _token_ranges(erm->get_primary_ranges(ep)) {}

				        std::size_t size() const { return _token_ranges.size(); }

				        const dht::token_range& operator[](std::size_t i) const {

				            return _token_ranges[i];

				        }

				        bool should_skip(std::size_t i) const {

				            return false;

				        }

				    };

				    // ranges_holder<secondary> holds the secondary token ranges plus each

				    // range's primary owner, needed to implement should_skip().

				    class ranges_holder_secondary {

				        std::vector<std::pair<dht::token_range, gms::inet_address>> _token_ranges;

				        gms::gossiper& _gossiper;

				     public:

				        ranges_holder_secondary(const locator::effective_replication_map_ptr& erm, gms::gossiper& g, gms::inet_address ep)

				            : _token_ranges(get_secondary_ranges(erm, ep))

				            , _gossiper(g) {}

				        std::size_t size() const { return _token_ranges.size(); }

				        const dht::token_range& operator[](std::size_t i) const {

				            return _token_ranges[i].first;

				        }

				        // range i should be skipped if its primary owner is alive.

				        bool should_skip(std::size_t i) const {

				            return _gossiper.is_alive(_token_ranges[i].second);

				        }

				    };

				// ranges_holder_primary holds just the primary ranges themselves

				class ranges_holder_primary {

				    dht::token_range_vector _token_ranges;

				public:

				    explicit ranges_holder_primary(dht::token_range_vector token_ranges) : _token_ranges(std::move(token_ranges)) {}

				    static future<ranges_holder_primary> make(const locator::vnode_effective_replication_map_ptr& erm, gms::inet_address ep) {

				        co_return ranges_holder_primary(co_await erm->get_primary_ranges(ep));

				    }

				    std::size_t size() const { return _token_ranges.size(); }

				    const dht::token_range& operator[](std::size_t i) const {

				        return _token_ranges[i];

				    }

				    bool should_skip(std::size_t i) const {

				        return false;

				    }

				};

				// ranges_holder<secondary> holds the secondary token ranges plus each

				// range's primary owner, needed to implement should_skip().

				class ranges_holder_secondary {

				    std::vector<std::pair<dht::token_range, gms::inet_address>> _token_ranges;

				    const gms::gossiper& _gossiper;

				public:

				    explicit ranges_holder_secondary(std::vector<std::pair<dht::token_range, gms::inet_address>> token_ranges, const gms::gossiper& g)

				        : _token_ranges(std::move(token_ranges))

				        , _gossiper(g) {}

				    static future<ranges_holder_secondary> make(const locator::effective_replication_map_ptr& erm, gms::inet_address ep, const gms::gossiper& g) {

				        co_return ranges_holder_secondary(co_await get_secondary_ranges(erm, ep), g);

				    }

				    std::size_t size() const { return _token_ranges.size(); }

				    const dht::token_range& operator[](std::size_t i) const {

				        return _token_ranges[i].first;

				    }

				    // range i should be skipped if its primary owner is alive.

				    bool should_skip(std::size_t i) const {

				        return _gossiper.is_alive(_token_ranges[i].second);

				    }

				};

				template<class primary_or_secondary_t>

				class token_ranges_owned_by_this_shard {

				    schema_ptr _s;

				    locator::effective_replication_map_ptr _erm;

				    // _token_ranges will contain a list of token ranges owned by this node.

				    // We'll further need to split each such range to the pieces owned by

				    // the current shard, using _intersecter.

				    using ranges_holder = std::conditional_t<

				            primary_or_secondary == primary_or_secondary_t::primary,

				            ranges_holder_primary,

				            ranges_holder_secondary>;

				    const ranges_holder _token_ranges;

				    const primary_or_secondary_t _token_ranges;

				    // NOTICE: _range_idx is used modulo _token_ranges size when accessing

				    // the data to ensure that it doesn't go out of bounds

				    size_t _range_idx;

				    size_t _end_idx;

				    std::optional<dht::selective_token_range_sharder> _intersecter;

				public:

				    token_ranges_owned_by_this_shard(replica::database& db, gms::gossiper& g, schema_ptr s)

				    token_ranges_owned_by_this_shard(schema_ptr s, primary_or_secondary_t token_ranges)

				        :  _s(s)

				        , _erm(s->table().get_effective_replication_map())

				        , _token_ranges(db.find_keyspace(s->ks_name()).get_vnode_effective_replication_map(),

				                g, _erm->get_topology().my_address())

				        , _token_ranges(std::move(token_ranges))

				        , _range_idx(random_offset(0, _token_ranges.size() - 1))

				        , _end_idx(_range_idx + _token_ranges.size())

				    {

				@@ -498,6 +502,7 @@ struct scan_ranges_context {

				    bytes column_name;

				    std::optional<std::string> member;

				    service::client_state internal_client_state;

				    ::shared_ptr<cql3::selection::selection> selection;

				    std::unique_ptr<service::query_state> query_state_ptr;

				    std::unique_ptr<cql3::query_options> query_options;

				@@ -507,6 +512,7 @@ struct scan_ranges_context {

				        : s(s)

				        , column_name(column_name)

				        , member(member)

				        , internal_client_state(service::client_state::internal_tag())

				    {

				        // FIXME: don't read the entire items - read only parts of it.

				        // We must read the key columns (to be able to delete) and also

				@@ -515,8 +521,9 @@ struct scan_ranges_context {

				        // be good if we can read only the single item of the map - it

				        // should be possible (and a must for issue #7751!).

				        lw_shared_ptr<service::pager::paging_state> paging_state = nullptr;

				        auto regular_columns = boost::copy_range<query::column_id_vector>(

				            s->regular_columns() | boost::adaptors::transformed([] (const column_definition& cdef) { return cdef.id; }));

				        auto regular_columns =

				            s->regular_columns() | std::views::transform([] (const column_definition& cdef) { return cdef.id; })

				            | std::ranges::to<query::column_id_vector>();

				        selection = cql3::selection::selection::wildcard(s);

				        query::partition_slice::option_set opts = selection->get_query_options();

				        opts.set<query::partition_slice::option::allow_short_read>();

				@@ -525,10 +532,9 @@ struct scan_ranges_context {

				        std::vector<query::clustering_range> ck_bounds{query::clustering_range::make_open_ended_both_sides()};

				        auto partition_slice = query::partition_slice(std::move(ck_bounds), {}, std::move(regular_columns), opts);

				        command = ::make_lw_shared<query::read_command>(s->id(), s->version(), partition_slice, proxy.get_max_result_size(partition_slice), query::tombstone_limit(proxy.get_tombstone_limit()));

				        executor::client_state client_state{executor::client_state::internal_tag()};

				        tracing::trace_state_ptr trace_state;

				        // NOTICE: empty_service_permit is used because the TTL service has fixed parallelism

				        query_state_ptr = std::make_unique<service::query_state>(client_state, trace_state, empty_service_permit());

				        query_state_ptr = std::make_unique<service::query_state>(internal_client_state, trace_state, empty_service_permit());

				        // FIXME: What should we do on multi-DC? Will we run the expiration on the same ranges on all

				        // DCs or only once for each range? If the latter, we need to change the CLs in the

				        // scanner and deleter.

				@@ -551,7 +557,7 @@ static future<> scan_table_ranges(

				        expiration_service::stats& expiration_stats)

				{

				    const schema_ptr& s = scan_ctx.s;

				    assert (partition_ranges.size() == 1); // otherwise issue #9167 will cause incorrect results.

				    SCYLLA_ASSERT (partition_ranges.size() == 1); // otherwise issue #9167 will cause incorrect results.

				    auto p = service::pager::query_pagers::pager(proxy, s, scan_ctx.selection, *scan_ctx.query_state_ptr,

				            *scan_ctx.query_options, scan_ctx.command, std::move(partition_ranges), nullptr);

				    while (!p->is_exhausted()) {

				@@ -724,7 +730,9 @@ static future<bool> scan_table(

				    expiration_stats.scan_table++;

				    // FIXME: need to pace the scan, not do it all at once.

				    scan_ranges_context scan_ctx{s, proxy, std::move(column_name), std::move(member)};

				    token_ranges_owned_by_this_shard<primary> my_ranges(db.real_database(), gossiper, s);

				    auto erm = db.real_database().find_keyspace(s->ks_name()).get_vnode_effective_replication_map();

				    auto my_address = erm->get_topology().my_address();

				    token_ranges_owned_by_this_shard my_ranges(s, co_await ranges_holder_primary::make(erm, my_address));

				    while (std::optional<dht::partition_range> range = my_ranges.next_partition_range()) {

				        // Note that because of issue #9167 we need to run a separate

				        // query on each partition range, and can't pass several of

				@@ -744,7 +752,7 @@ static future<bool> scan_table(

				    // by tasking another node to take over scanning of the dead node's primary

				    // ranges. What we do here is that this node will also check expiration

				    // on its *secondary* ranges - but only those whose primary owner is down.

				    token_ranges_owned_by_this_shard<secondary> my_secondary_ranges(db.real_database(), gossiper, s);

				    token_ranges_owned_by_this_shard my_secondary_ranges(s, co_await ranges_holder_secondary::make(erm, my_address, gossiper));

				    while (std::optional<dht::partition_range> range = my_secondary_ranges.next_partition_range()) {

				        expiration_stats.secondary_ranges_scanned++;

				        dht::partition_range_vector partition_ranges;

									
										31

api/CMakeLists.txt
									
												View File
												
				@@ -1,4 +1,29 @@

				# Generate C++ sources from Swagger definitions

				function(generate_swagger)

				  set(one_value_args TARGET VAR IN_FILE OUT_DIR)

				  cmake_parse_arguments(args "" "${one_value_args}" "" ${ARGN})

				  get_filename_component(in_file_name ${args_IN_FILE} NAME)

				  set(generator ${PROJECT_SOURCE_DIR}/seastar/scripts/seastar-json2code.py)

				  set(header_out ${args_OUT_DIR}/${in_file_name}.hh)

				  set(source_out ${args_OUT_DIR}/${in_file_name}.cc)

				  add_custom_command(

				    DEPENDS

				      ${args_IN_FILE}

				      ${generator}

				    OUTPUT ${header_out} ${source_out}

				    COMMAND ${CMAKE_COMMAND} -E make_directory ${args_OUT_DIR}

				    COMMAND ${generator} --create-cc -f ${args_IN_FILE} -o ${header_out})

				  add_custom_target(${args_TARGET}

				    DEPENDS

				      ${header_out}

				      ${source_out})

				  set(${args_VAR} ${header_out} ${source_out} PARENT_SCOPE)

				endfunction()

				set(swagger_files

				  api-doc/authorization_cache.json

				  api-doc/cache_service.json

				@@ -7,6 +32,7 @@ set(swagger_files

				  api-doc/commitlog.json

				  api-doc/compaction_manager.json

				  api-doc/config.json

				  api-doc/cql_server_test.json

				  api-doc/endpoint_snitch_info.json

				  api-doc/error_injection.json

				  api-doc/failure_detector.json

				@@ -28,7 +54,7 @@ set(swagger_files

				foreach(f ${swagger_files})

				  get_filename_component(fname "${f}" NAME_WE)

				  get_filename_component(dir "${f}" DIRECTORY)

				  seastar_generate_swagger(

				  generate_swagger(

				    TARGET scylla_swagger_gen_${fname}

				    VAR scylla_swagger_gen_${fname}_files

				    IN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/${f}"

				@@ -36,7 +62,7 @@ foreach(f ${swagger_files})

				  list(APPEND swagger_gen_files "${scylla_swagger_gen_${fname}_files}")

				endforeach()

				add_library(api)

				add_library(api STATIC)

				target_sources(api

				  PRIVATE

				    api.cc

				@@ -46,6 +72,7 @@ target_sources(api

				    commitlog.cc

				    compaction_manager.cc

				    config.cc

				    cql_server_test.cc

				    endpoint_snitch.cc

				    error_injection.cc

				    authorization_cache.cc

									
										8

api/api-doc/column_family.json
									
												View File
												
				@@ -92,6 +92,14 @@

				                     "type":"boolean",

				                     "paramType":"query"

				                  },

				                  {

				                     "name":"consider_only_existing_data",

				                     "description":"Set to \"true\" to flush all memtables and force tombstone garbage collection to check only the sstables being compacted (false by default). The memtable, commitlog and other uncompacted sstables will not be checked during tombstone garbage collection.",

				                     "required":false,

				                     "allowMultiple":false,

				                     "type":"boolean",

				                     "paramType":"query"

				                  },

				                  {

				                     "name":"split_output",

				                     "description":"true if the output of the major compaction should be split in several sstables",

									
										26

api/api-doc/cql_server_test.json
									
										Normal file
									
												View File
												
				@@ -0,0 +1,26 @@

				{

				    "apiVersion":"0.0.1",

				    "swaggerVersion":"1.2",

				    "basePath":"{{Protocol}}://{{Host}}",

				    "resourcePath":"/cql_server_test",

				    "produces":[

				        "application/json"

				    ],

				    "apis":[

				        {

				            "path":"/cql_server_test/connections_params",

				            "operations":[

				                {

				                    "method":"GET",

				                    "summary":"Get service level params of each CQL connection",

				                    "type":"connections_service_level_params",

				                    "nickname":"connections_params",

				                    "produces":[

				                        "application/json"

				                    ],

				                    "parameters":[]

				                }

				            ]

				        }

				    ]

				}

									
										32

api/api-doc/raft.json
									
												View File
												
				@@ -94,6 +94,38 @@

				               ]

				            }

				         ]

				      },

				      {

				         "path":"/raft/trigger_stepdown/",

				         "operations":[

				            {

				               "method":"POST",

				               "summary":"Triggers stepdown of a leader for given Raft group or group0 if not provided (returns an error if the node is not a leader)",

				               "type":"string",

				               "nickname":"trigger_stepdown",

				               "produces":[

				                  "application/json"

				               ],

				               "parameters":[

				                  {

				                     "name":"group_id",

				                     "description":"The ID of the group which leader should stepdown",

				                     "required":false,

				                     "allowMultiple":false,

				                     "type":"string",

				                     "paramType":"query"

				                  },

				                  {

				                     "name":"timeout",

				                     "description":"Timeout in seconds after which the endpoint returns a failure. If not provided, 60s is used.",

				                     "required":false,

				                     "allowMultiple":false,

				                     "type":"long",

				                     "paramType":"query"

				                  }

				               ]

				            }

				         ]

				      }

				   ]

				}

									
										156

api/api-doc/storage_service.json
									
												View File
												
				@@ -741,11 +741,151 @@

				                     "allowMultiple":false,

				                     "type":"boolean",

				                     "paramType":"query"

				                  },

				                  {

				                     "name":"consider_only_existing_data",

				                     "description":"Set to \"true\" to flush all memtables and force tombstone garbage collection to check only the sstables being compacted (false by default). The memtable, commitlog and other uncompacted sstables will not be checked during tombstone garbage collection.",

				                     "required":false,

				                     "allowMultiple":false,

				                     "type":"boolean",

				                     "paramType":"query"

				                  }

				               ]

				            }

				         ]

				      },

				      {

				          "path":"/storage_service/backup",

				          "operations":[

				              {

				                  "method":"POST",

				                  "summary":"Starts copying SSTables from a specified keyspace to a designated bucket in object storage",

				                  "type":"string",

				                  "nickname":"start_backup",

				                  "produces":[

				                      "application/json"

				                  ],

				                  "parameters":[

				                      {

				                          "name":"endpoint",

				                          "description":"ID of the configured object storage endpoint to copy sstables to",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                          "name":"bucket",

				                          "description":"Name of the bucket to backup sstables to",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                           "name":"prefix",

				                           "description":"The prefix of the objects for the backuped sstables",

				                           "required":true,

				                           "allowMultiple":false,

				                           "type":"string",

				                           "paramType":"query"

				                      },

				                      {

				                          "name":"keyspace",

				                          "description":"Name of a keyspace to copy sstables from",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                          "name":"table",

				                          "description":"Name of a table to copy sstables from",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                          "name":"snapshot",

				                          "description":"Name of a snapshot to copy sstables from",

				                          "required":false,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      }

				                  ]

				              }

				          ]

				      },

				      {

				          "path":"/storage_service/restore",

				          "operations":[

				              {

				                  "method":"POST",

				                  "summary":"Starts copying SSTables from a designated bucket in object storage to a specified keyspace",

				                  "type":"string",

				                  "nickname":"start_restore",

				                  "produces":[

				                      "application/json"

				                  ],

				                  "parameters":[

				                      {

				                          "name":"endpoint",

				                          "description":"ID of the configured object storage endpoint to copy SSTables from",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                          "name":"bucket",

				                          "description":"Name of the bucket to read SSTables from",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                          "name":"prefix",

				                          "description":"The prefix of the object keys for the backuped SSTables",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                          "in": "body",

				                          "name": "sstables",

				                          "description": "The list of the object keys of the TOC component of the SSTables to be restored",

				                          "required":true,

				                          "schema" :{

				                              "type": "array",

				                              "items": {

				                                  "type": "string"

				                              }

				                          }

				                      },

				                      {

				                          "name":"keyspace",

				                          "description":"Name of a keyspace to copy SSTables to",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      },

				                      {

				                          "name":"table",

				                          "description":"Name of a table to copy SSTables to",

				                          "required":true,

				                          "allowMultiple":false,

				                          "type":"string",

				                          "paramType":"query"

				                      }

				                  ]

				              }

				          ]

				      },

				      {

				         "path":"/storage_service/keyspace_compaction/{keyspace}",

				         "operations":[

				@@ -781,6 +921,14 @@

				                     "allowMultiple":false,

				                     "type":"boolean",

				                     "paramType":"query"

				                  },

				                  {

				                     "name":"consider_only_existing_data",

				                     "description":"Set to \"true\" to flush all memtables and force tombstone garbage collection to check only the sstables being compacted (false by default). The memtable, commitlog and other uncompacted sstables will not be checked during tombstone garbage collection.",

				                     "required":false,

				                     "allowMultiple":false,

				                     "type":"boolean",

				                     "paramType":"query"

				                  }

				               ]

				            }

				@@ -1891,6 +2039,14 @@

				                     "allowMultiple":false,

				                     "type":"string",

				                     "paramType":"query"

				                  },

				                  {

				                     "name":"force",

				                     "description":"Enforce the source_dc option, even if it unsafe to use for rebuild",

				                     "required":false,

				                     "allowMultiple":false,

				                     "type":"boolean",

				                     "paramType":"query"

				                  }

				               ]

				            }

									
										15

api/api-doc/system.json
									
												View File
												
				@@ -194,6 +194,21 @@

				               "parameters":[]

				            }

				         ]

				      },

				      {

				         "path":"/system/highest_supported_sstable_version",

				         "operations":[

				            {

				               "method":"GET",

				               "summary":"Get highest supported sstable version",

				               "type":"string",

				               "nickname":"get_highest_supported_sstable_version",

				               "produces":[

				                  "application/json"

				               ],

				               "parameters":[]

				            }

				         ]

				      }

				   ]

				}

									
										55

api/api-doc/task_manager.json
									
												View File
												
				@@ -115,7 +115,7 @@

				               "parameters":[

				                  {

				                     "name":"task_id",

				                     "description":"The uuid of a task to abort",

				                     "description":"The uuid of a task to abort; if the task is not abortable, 403 status code is returned",

				                     "required":true,

				                     "allowMultiple":false,

				                     "type":"string",

				@@ -144,6 +144,14 @@

				                     "allowMultiple":false,

				                     "type":"string",

				                     "paramType":"path"

				                  },

				                  {

				                     "name":"timeout",

				                     "description":"Timeout for waiting; if times out, 408 status code is returned",

				                     "required":false,

				                     "allowMultiple":false,

				                     "type":"long",

				                     "paramType":"query"

				                  }

				               ]

				            }

				@@ -197,11 +205,36 @@

				                     "paramType":"query"

				                  }

				               ]

				            },

				            {

				               "method":"GET",

				               "summary":"Get current ttl value",

				               "type":"long",

				               "nickname":"get_ttl",

				               "produces":[

				                  "application/json"

				               ],

				               "parameters":[

				               ]

				            }

				         ]

				      }

				   ],

				   "models":{

				      "task_identity":{

				         "id": "task_identity",

				         "description":"Id and node of a task",

				         "properties":{

				            "task_id":{

				               "type":"string",

				               "description":"The uuid of a task"

				            },

				            "node":{

				               "type":"string",

				               "description":"Address of a server on which a task is created"

				            }

				         }

				      },

				      "task_stats" :{

				         "id": "task_stats",

				         "description":"A task statistics object",

				@@ -224,6 +257,14 @@

				               "type":"string",

				               "description":"The description of the task"

				            },

				            "kind":{

				               "type":"string",

				               "enum":[

				                  "node",

				                  "cluster"

				               ],

				               "description":"The kind of a task"

				            },

				            "scope":{

				               "type":"string",

				               "description":"The scope of the task"

				@@ -258,6 +299,14 @@

				               "type":"string",

				               "description":"The description of the task"

				            },

				            "kind":{

				               "type":"string",

				               "enum":[

				                  "node",

				                  "cluster"

				               ],

				               "description":"The kind of a task"

				            },

				            "scope":{

				               "type":"string",

				               "description":"The scope of the task"

				@@ -327,9 +376,9 @@

				            "children_ids":{

				               "type":"array",

				               "items":{

				                  "type":"string"

				                  "type":"task_identity"

				               },

				               "description":"Task IDs of children of this task"

				               "description":"Task identities of children of this task"

				            }

				         }

				      }

									
										54

api/api.cc
									
												View File
												
				@@ -10,6 +10,7 @@

				#include <seastar/http/file_handler.hh>

				#include <seastar/http/transformers.hh>

				#include <seastar/http/api_docs.hh>

				#include "cql_server_test.hh"

				#include "storage_service.hh"

				#include "token_metadata.hh"

				#include "commitlog.hh"

				@@ -73,6 +74,8 @@ future<> set_server_init(http_context& ctx) {

				        set_error_injection(ctx, r);

				        rb->register_function(r, "storage_proxy",

				                "The storage proxy API");

				        rb->register_function(r, "storage_service",

				                "The storage service API");

				    });

				}

				@@ -115,7 +118,7 @@ future<> unset_thrift_controller(http_context& ctx) {

				}

				future<> set_server_storage_service(http_context& ctx, sharded<service::storage_service>& ss, service::raft_group0_client& group0_client) {

				    return register_api(ctx, "storage_service", "The storage service API", [&ss, &group0_client] (http_context& ctx, routes& r) {

				    return ctx.http_server.set_routes([&ctx, &ss, &group0_client] (routes& r) {

				            set_storage_service(ctx, r, ss, group0_client);

				        });

				}

				@@ -132,6 +135,14 @@ future<> unset_load_meter(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_load_meter(ctx, r); });

				}

				future<> set_format_selector(http_context& ctx, db::sstables_format_selector& sel) {

				    return ctx.http_server.set_routes([&ctx, &sel] (routes& r) { set_format_selector(ctx, r, sel); });

				}

				future<> unset_format_selector(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_format_selector(ctx, r); });

				}

				future<> set_server_sstables_loader(http_context& ctx, sharded<sstables_loader>& sst_loader) {

				    return ctx.http_server.set_routes([&ctx, &sst_loader] (routes& r) { set_sstables_loader(ctx, r, sst_loader); });

				}

				@@ -256,6 +267,10 @@ future<> set_server_cache(http_context& ctx) {

				            "The cache service API", set_cache_service);

				}

				future<> unset_server_cache(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_cache_service(ctx, r); });

				}

				future<> set_hinted_handoff(http_context& ctx, sharded<service::storage_proxy>& proxy) {

				    return register_api(ctx, "hinted_handoff",

				                "The hinted handoff API", [&proxy] (http_context& ctx, routes& r) {

				@@ -267,16 +282,16 @@ future<> unset_hinted_handoff(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_hinted_handoff(ctx, r); });

				}

				future<> set_server_compaction_manager(http_context& ctx) {

				    auto rb = std::make_shared < api_registry_builder > (ctx.api_doc);

				    return ctx.http_server.set_routes([rb, &ctx](routes& r) {

				        rb->register_function(r, "compaction_manager",

				                "The Compaction manager API");

				        set_compaction_manager(ctx, r);

				future<> set_server_compaction_manager(http_context& ctx, sharded<compaction_manager>& cm) {

				    return register_api(ctx, "compaction_manager", "The Compaction manager API", [&cm] (http_context& ctx, routes& r) {

				        set_compaction_manager(ctx, r, cm);

				    });

				}

				future<> unset_server_compaction_manager(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_compaction_manager(ctx, r); });

				}

				future<> set_server_done(http_context& ctx) {

				    auto rb = std::make_shared < api_registry_builder > (ctx.api_doc);

				@@ -284,15 +299,22 @@ future<> set_server_done(http_context& ctx) {

				        rb->register_function(r, "lsa", "Log-structured allocator API");

				        set_lsa(ctx, r);

				        rb->register_function(r, "commitlog",

				                "The commit log API");

				        set_commitlog(ctx,r);

				        rb->register_function(r, "collectd",

				                "The collectd API");

				        set_collectd(ctx, r);

				    });

				}

				future<> set_server_commitlog(http_context& ctx, sharded<replica::database>& db) {

				    return register_api(ctx, "commitlog", "The commit log API", [&db] (http_context& ctx, routes& r) {

				        set_commitlog(ctx, r, db);

				    });

				}

				future<> unset_server_commitlog(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_commitlog(ctx, r); });

				}

				future<> set_server_task_manager(http_context& ctx, sharded<tasks::task_manager>& tm, lw_shared_ptr<db::config> cfg) {

				    auto rb = std::make_shared < api_registry_builder > (ctx.api_doc);

				@@ -323,6 +345,16 @@ future<> unset_server_task_manager_test(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_task_manager_test(ctx, r); });

				}

				future<> set_server_cql_server_test(http_context& ctx, cql_transport::controller& ctl) {

				    return register_api(ctx, "cql_server_test", "The CQL server test API", [&ctl] (http_context& ctx, routes& r) {

				        set_cql_server_test(ctx, r, ctl);

				    });

				}

				future<> unset_server_cql_server_test(http_context& ctx) {

				    return ctx.http_server.set_routes([&ctx] (routes& r) { unset_cql_server_test(ctx, r); });

				}

				#endif

				future<> set_server_tasks_compaction_module(http_context& ctx, sharded<service::storage_service>& ss, sharded<db::snapshot_ctl>& snap_ctl) {

									
										2

api/api.hh
									
												View File
												
				@@ -246,7 +246,7 @@ public:

				                value = T{boost::lexical_cast<Base>(param)};

				            }

				        } catch (boost::bad_lexical_cast&) {

				            throw httpd::bad_param_exception(format("{} ({}): type error - should be {}", name, param, boost::units::detail::demangle(typeid(Base).name())));

				            throw httpd::bad_param_exception(fmt::format("{} ({}): type error - should be {}", name, param, boost::units::detail::demangle(typeid(Base).name())));

				        }

				    }

									
										13

api/api_init.hh
									
												View File
												
				@@ -17,6 +17,8 @@

				using request = http::request;

				using reply = http::reply;

				class compaction_manager;

				namespace service {

				class load_meter;

				@@ -49,6 +51,7 @@ namespace cql_transport { class controller; }

				namespace db {

				class snapshot_ctl;

				class config;

				class sstables_format_selector;

				namespace view {

				class view_builder;

				}

				@@ -119,7 +122,9 @@ future<> unset_server_stream_manager(http_context& ctx);

				future<> set_hinted_handoff(http_context& ctx, sharded<service::storage_proxy>& p);

				future<> unset_hinted_handoff(http_context& ctx);

				future<> set_server_cache(http_context& ctx);

				future<> set_server_compaction_manager(http_context& ctx);

				future<> unset_server_cache(http_context& ctx);

				future<> set_server_compaction_manager(http_context& ctx, sharded<compaction_manager>& cm);

				future<> unset_server_compaction_manager(http_context& ctx);

				future<> set_server_done(http_context& ctx);

				future<> set_server_task_manager(http_context& ctx, sharded<tasks::task_manager>& tm, lw_shared_ptr<db::config> cfg);

				future<> unset_server_task_manager(http_context& ctx);

				@@ -131,5 +136,11 @@ future<> set_server_raft(http_context&, sharded<service::raft_group_registry>&);

				future<> unset_server_raft(http_context&);

				future<> set_load_meter(http_context& ctx, service::load_meter& lm);

				future<> unset_load_meter(http_context& ctx);

				future<> set_format_selector(http_context& ctx, db::sstables_format_selector& sel);

				future<> unset_format_selector(http_context& ctx);

				future<> set_server_cql_server_test(http_context& ctx, cql_transport::controller& ctl);

				future<> unset_server_cql_server_test(http_context& ctx);

				future<> set_server_commitlog(http_context& ctx, sharded<replica::database>&);

				future<> unset_server_commitlog(http_context& ctx);

				}

									
										45

api/cache_service.cc
									
												View File
												
				@@ -320,5 +320,50 @@ void set_cache_service(http_context& ctx, routes& r) {

				    });

				}

				void unset_cache_service(http_context& ctx, routes& r) {

				    cs::get_row_cache_save_period_in_seconds.unset(r);

				    cs::set_row_cache_save_period_in_seconds.unset(r);

				    cs::get_key_cache_save_period_in_seconds.unset(r);

				    cs::set_key_cache_save_period_in_seconds.unset(r);

				    cs::get_counter_cache_save_period_in_seconds.unset(r);

				    cs::set_counter_cache_save_period_in_seconds.unset(r);

				    cs::get_row_cache_keys_to_save.unset(r);

				    cs::set_row_cache_keys_to_save.unset(r);

				    cs::get_key_cache_keys_to_save.unset(r);

				    cs::set_key_cache_keys_to_save.unset(r);

				    cs::get_counter_cache_keys_to_save.unset(r);

				    cs::set_counter_cache_keys_to_save.unset(r);

				    cs::invalidate_key_cache.unset(r);

				    cs::invalidate_counter_cache.unset(r);

				    cs::set_row_cache_capacity_in_mb.unset(r);

				    cs::set_key_cache_capacity_in_mb.unset(r);

				    cs::set_counter_cache_capacity_in_mb.unset(r);

				    cs::save_caches.unset(r);

				    cs::get_key_capacity.unset(r);

				    cs::get_key_hits.unset(r);

				    cs::get_key_requests.unset(r);

				    cs::get_key_hit_rate.unset(r);

				    cs::get_key_hits_moving_avrage.unset(r);

				    cs::get_key_requests_moving_avrage.unset(r);

				    cs::get_key_size.unset(r);

				    cs::get_key_entries.unset(r);

				    cs::get_row_capacity.unset(r);

				    cs::get_row_hits.unset(r);

				    cs::get_row_requests.unset(r);

				    cs::get_row_hit_rate.unset(r);

				    cs::get_row_hits_moving_avrage.unset(r);

				    cs::get_row_requests_moving_avrage.unset(r);

				    cs::get_row_size.unset(r);

				    cs::get_row_entries.unset(r);

				    cs::get_counter_capacity.unset(r);

				    cs::get_counter_hits.unset(r);

				    cs::get_counter_requests.unset(r);

				    cs::get_counter_hit_rate.unset(r);

				    cs::get_counter_hits_moving_avrage.unset(r);

				    cs::get_counter_requests_moving_avrage.unset(r);

				    cs::get_counter_size.unset(r);

				    cs::get_counter_entries.unset(r);

				}

				}

									
										1

api/cache_service.hh
									
												View File
												
				@@ -16,5 +16,6 @@ namespace api {

				struct http_context;

				void set_cache_service(http_context& ctx, seastar::httpd::routes& r);

				void unset_cache_service(http_context& ctx, seastar::httpd::routes& r);

				}

									
										3

api/collectd.cc
									
												View File
												
				@@ -11,6 +11,7 @@

				#include <seastar/core/scollectd.hh>

				#include <seastar/core/scollectd_api.hh>

				#include <boost/range/irange.hpp>

				#include <ranges>

				#include <regex>

				#include "api/api_init.hh"

				@@ -61,7 +62,7 @@ void set_collectd(http_context& ctx, routes& r) {

				        return do_with(std::vector<cd::collectd_value>(), [id] (auto& vec) {

				            vec.resize(smp::count);

				            return parallel_for_each(boost::irange(0u, smp::count), [&vec, id] (auto cpu) {

				            return parallel_for_each(std::views::iota(0u, smp::count), [&vec, id] (auto cpu) {

				                return smp::submit_to(cpu, [id = *id] {

				                    return scollectd::get_collectd_value(id);

				                }).then([&vec, cpu] (auto res) {

									
										35

api/column_family.cc
									
												View File
												
				@@ -15,6 +15,7 @@

				#include <seastar/http/exception.hh>

				#include "sstables/sstables.hh"

				#include "sstables/metadata_collector.hh"

				#include "utils/assert.hh"

				#include "utils/estimated_histogram.hh"

				#include <algorithm>

				#include "db/system_keyspace.hh"

				@@ -23,6 +24,8 @@

				#include "compaction/compaction_manager.hh"

				#include "unimplemented.hh"

				#include <boost/range/algorithm/copy.hpp>

				extern logging::logger apilog;

				namespace api {

				@@ -60,14 +63,6 @@ table_id get_uuid(const sstring& name, const replica::database& db) {

				    return get_uuid(ks, cf, db);

				}

				future<> foreach_column_family(http_context& ctx, const sstring& name, std::function<void(replica::column_family&)> f) {

				    auto uuid = get_uuid(name, ctx.db.local());

				    return ctx.db.invoke_on_all([f, uuid](replica::database& db) {

				        f(db.find_column_family(uuid));

				    });

				}

				future<json::json_return_type>  get_cf_stats(http_context& ctx, const sstring& name,

				        int64_t replica::column_family_stats::*f) {

				    return map_reduce_cf(ctx, name, int64_t(0), [f](const replica::column_family& cf) {

				@@ -82,7 +77,7 @@ future<json::json_return_type>  get_cf_stats(http_context& ctx,

				    }, std::plus<int64_t>());

				}

				static future<json::json_return_type> set_tables(http_context& ctx, const sstring& keyspace, std::vector<sstring> tables, std::function<future<>(replica::table&)> set) {

				static future<json::json_return_type> for_tables_on_all_shards(http_context& ctx, const sstring& keyspace, std::vector<sstring> tables, std::function<future<>(replica::table&)> set) {

				    if (tables.empty()) {

				        tables = map_keys(ctx.db.local().find_keyspace(keyspace).metadata().get()->cf_meta_data());

				    }

				@@ -103,7 +98,7 @@ class autocompaction_toggle_guard {

				    replica::database& _db;

				public:

				    autocompaction_toggle_guard(replica::database& db) : _db(db) {

				        assert(this_shard_id() == 0);

				        SCYLLA_ASSERT(this_shard_id() == 0);

				        if (!_db._enable_autocompaction_toggle) {

				            throw std::runtime_error("Autocompaction toggle is busy");

				        }

				@@ -112,7 +107,7 @@ public:

				    autocompaction_toggle_guard(const autocompaction_toggle_guard&) = delete;

				    autocompaction_toggle_guard(autocompaction_toggle_guard&&) = default;

				    ~autocompaction_toggle_guard() {

				        assert(this_shard_id() == 0);

				        SCYLLA_ASSERT(this_shard_id() == 0);

				        _db._enable_autocompaction_toggle = true;

				    }

				};

				@@ -122,7 +117,7 @@ static future<json::json_return_type> set_tables_autocompaction(http_context& ct

				    return ctx.db.invoke_on(0, [&ctx, keyspace, tables = std::move(tables), enabled] (replica::database& db) {

				        auto g = autocompaction_toggle_guard(db);

				        return set_tables(ctx, keyspace, tables, [enabled] (replica::table& cf) {

				        return for_tables_on_all_shards(ctx, keyspace, tables, [enabled] (replica::table& cf) {

				            if (enabled) {

				                cf.enable_auto_compaction();

				            } else {

				@@ -135,7 +130,7 @@ static future<json::json_return_type> set_tables_autocompaction(http_context& ct

				static future<json::json_return_type> set_tables_tombstone_gc(http_context& ctx, const sstring &keyspace, std::vector<sstring> tables, bool enabled) {

				    apilog.info("set_tables_tombstone_gc: enabled={} keyspace={} tables={}", enabled, keyspace, tables);

				    return set_tables(ctx, keyspace, std::move(tables), [enabled] (replica::table& t) {

				    return for_tables_on_all_shards(ctx, keyspace, std::move(tables), [enabled] (replica::table& t) {

				        t.set_tombstone_gc_enabled(enabled);

				        return make_ready_future<>();

				    });

				@@ -1055,12 +1050,12 @@ void set_column_family(http_context& ctx, routes& r, sharded<db::system_keyspace

				    });

				    cf::set_compaction_strategy_class.set(r, [&ctx](std::unique_ptr<http::request> req) {

				        auto [ks, cf] = parse_fully_qualified_cf_name(req->get_path_param("name"));

				        sstring strategy = req->get_query_param("class_name");

				        apilog.info("column_family/set_compaction_strategy_class: name={} strategy={}", req->get_path_param("name"), strategy);

				        return foreach_column_family(ctx, req->get_path_param("name"), [strategy](replica::column_family& cf) {

				        return for_tables_on_all_shards(ctx, ks, {std::move(cf)}, [strategy] (replica::table& cf) {

				            cf.set_compaction_strategy(sstables::compaction_strategy::type(strategy));

				        }).then([] {

				                return make_ready_future<json::json_return_type>(json_void());

				            return make_ready_future<>();

				        });

				    });

				@@ -1125,6 +1120,7 @@ void set_column_family(http_context& ctx, routes& r, sharded<db::system_keyspace

				        auto params = req_params({

				            std::pair("name", mandatory::yes),

				            std::pair("flush_memtables", mandatory::no),

				            std::pair("consider_only_existing_data", mandatory::no),

				            std::pair("split_output", mandatory::no),

				        });

				        params.process(*req);

				@@ -1133,7 +1129,8 @@ void set_column_family(http_context& ctx, routes& r, sharded<db::system_keyspace

				        }

				        auto [ks, cf] = parse_fully_qualified_cf_name(*params.get("name"));

				        auto flush = params.get_as<bool>("flush_memtables").value_or(true);

				        apilog.info("column_family/force_major_compaction: name={} flush={}", req->get_path_param("name"), flush);

				        auto consider_only_existing_data = params.get_as<bool>("consider_only_existing_data").value_or(false);

				        apilog.info("column_family/force_major_compaction: name={} flush={} consider_only_existing_data={}", req->get_path_param("name"), flush, consider_only_existing_data);

				        auto keyspace = validate_keyspace(ctx, ks);

				        std::vector<table_info> table_infos = {table_info{

				@@ -1143,10 +1140,10 @@ void set_column_family(http_context& ctx, routes& r, sharded<db::system_keyspace

				        auto& compaction_module = ctx.db.local().get_compaction_manager().get_task_manager_module();

				        std::optional<flush_mode> fmopt;

				        if (!flush) {

				        if (!flush && !consider_only_existing_data) {

				            fmopt = flush_mode::skip;

				        }

				        auto task = co_await compaction_module.make_and_start_task<major_keyspace_compaction_task_impl>({}, std::move(keyspace), tasks::task_id::create_null_id(), ctx.db, std::move(table_infos), fmopt);

				        auto task = co_await compaction_module.make_and_start_task<major_keyspace_compaction_task_impl>({}, std::move(keyspace), tasks::task_id::create_null_id(), ctx.db, std::move(table_infos), fmopt, consider_only_existing_data);

				        co_await task->done();

				        co_return json_void();

				    });

									
										2

api/column_family.hh
									
												View File
												
				@@ -23,8 +23,6 @@ void set_column_family(http_context& ctx, httpd::routes& r, sharded<db::system_k

				void unset_column_family(http_context& ctx, httpd::routes& r);

				table_id get_uuid(const sstring& name, const replica::database& db);

				future<> foreach_column_family(http_context& ctx, const sstring& name, std::function<void(replica::column_family&)> f);

				template<class Mapper, class I, class Reducer>

				future<I> map_reduce_cf_raw(http_context& ctx, const sstring& name, I init,

									
										43

api/commitlog.cc
									
												View File
												
				@@ -9,18 +9,20 @@

				#include "commitlog.hh"

				#include "db/commitlog/commitlog.hh"

				#include "api/api-doc/commitlog.json.hh"

				#include "api/api-doc/storage_service.json.hh"

				#include "api/api_init.hh"

				#include "replica/database.hh"

				#include <vector>

				namespace api {

				using namespace seastar::httpd;

				namespace ss = httpd::storage_service_json;

				template<typename T>

				static auto acquire_cl_metric(http_context& ctx, std::function<T (const db::commitlog*)> func) {

				static auto acquire_cl_metric(sharded<replica::database>& db, std::function<T (const db::commitlog*)> func) {

				    typedef T ret_type;

				    return ctx.db.map_reduce0([func = std::move(func)](replica::database& db) {

				    return db.map_reduce0([func = std::move(func)](replica::database& db) {

				        if (db.commitlog() == nullptr) {

				            return make_ready_future<ret_type>();

				        }

				@@ -30,11 +32,11 @@ static auto acquire_cl_metric(http_context& ctx, std::function<T (const db::comm

				    });

				}

				void set_commitlog(http_context& ctx, routes& r) {

				void set_commitlog(http_context& ctx, routes& r, sharded<replica::database>& db) {

				    httpd::commitlog_json::get_active_segment_names.set(r,

				            [&ctx](std::unique_ptr<request> req) {

				            [&db](std::unique_ptr<request> req) {

				        auto res = make_shared<std::vector<sstring>>();

				        return ctx.db.map_reduce([res](std::vector<sstring> names) {

				        return db.map_reduce([res](std::vector<sstring> names) {

				            res->insert(res->end(), names.begin(), names.end());

				        }, [](replica::database& db) {

				            if (db.commitlog() == nullptr) {

				@@ -52,20 +54,35 @@ void set_commitlog(http_context& ctx, routes& r) {

				        return res;

				    });

				    httpd::commitlog_json::get_completed_tasks.set(r, [&ctx](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(ctx, std::bind(&db::commitlog::get_completed_tasks, std::placeholders::_1));

				    httpd::commitlog_json::get_completed_tasks.set(r, [&db](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(db, std::bind(&db::commitlog::get_completed_tasks, std::placeholders::_1));

				    });

				    httpd::commitlog_json::get_pending_tasks.set(r, [&ctx](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(ctx, std::bind(&db::commitlog::get_pending_tasks, std::placeholders::_1));

				    httpd::commitlog_json::get_pending_tasks.set(r, [&db](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(db, std::bind(&db::commitlog::get_pending_tasks, std::placeholders::_1));

				    });

				    httpd::commitlog_json::get_total_commit_log_size.set(r, [&ctx](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(ctx, std::bind(&db::commitlog::get_total_size, std::placeholders::_1));

				    httpd::commitlog_json::get_total_commit_log_size.set(r, [&db](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(db, std::bind(&db::commitlog::get_total_size, std::placeholders::_1));

				    });

				    httpd::commitlog_json::get_max_disk_size.set(r, [&ctx](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(ctx, std::bind(&db::commitlog::disk_limit, std::placeholders::_1));

				    httpd::commitlog_json::get_max_disk_size.set(r, [&db](std::unique_ptr<request> req) {

				        return acquire_cl_metric<uint64_t>(db, std::bind(&db::commitlog::disk_limit, std::placeholders::_1));

				    });

				    ss::get_commitlog.set(r, [&db](const_req req) {

				        return db.local().commitlog()->active_config().commit_log_location;

				    });

				}

				void unset_commitlog(http_context& ctx, routes& r) {

				    httpd::commitlog_json::get_active_segment_names.unset(r);

				    httpd::commitlog_json::get_archiving_segment_names.unset(r);

				    httpd::commitlog_json::get_completed_tasks.unset(r);

				    httpd::commitlog_json::get_pending_tasks.unset(r);

				    httpd::commitlog_json::get_total_commit_log_size.unset(r);

				    httpd::commitlog_json::get_max_disk_size.unset(r);

				    ss::get_commitlog.unset(r);

				}

				}

									
										7

api/commitlog.hh
									
												View File
												
				@@ -8,12 +8,17 @@

				#pragma once

				#include <seastar/core/sharded.hh>

				namespace seastar::httpd {

				class routes;

				}

				namespace replica { class database; }

				namespace api {

				struct http_context;

				void set_commitlog(http_context& ctx, seastar::httpd::routes& r);

				void set_commitlog(http_context& ctx, seastar::httpd::routes& r, seastar::sharded<replica::database>&);

				void unset_commitlog(http_context& ctx, seastar::httpd::routes& r);

				}

									
										62

api/compaction_manager.cc
									
												View File
												
				@@ -13,6 +13,7 @@

				#include "compaction/compaction_manager.hh"

				#include "api/api.hh"

				#include "api/api-doc/compaction_manager.json.hh"

				#include "api/api-doc/storage_service.json.hh"

				#include "db/system_keyspace.hh"

				#include "column_family.hh"

				#include "unimplemented.hh"

				@@ -23,13 +24,14 @@

				namespace api {

				namespace cm = httpd::compaction_manager_json;

				namespace ss = httpd::storage_service_json;

				using namespace json;

				using namespace seastar::httpd;

				static future<json::json_return_type> get_cm_stats(http_context& ctx,

				static future<json::json_return_type> get_cm_stats(sharded<compaction_manager>& cm,

				        int64_t compaction_manager::stats::*f) {

				    return ctx.db.map_reduce0([f](replica::database& db) {

				        return db.get_compaction_manager().get_stats().*f;

				    return cm.map_reduce0([f](compaction_manager& cm) {

				        return cm.get_stats().*f;

				    }, int64_t(0), std::plus<int64_t>()).then([](const int64_t& res) {

				        return make_ready_future<json::json_return_type>(res);

				    });

				@@ -44,11 +46,10 @@ static std::unordered_map<std::pair<sstring, sstring>, uint64_t, utils::tuple_ha

				    return std::move(a);

				}

				void set_compaction_manager(http_context& ctx, routes& r) {

				    cm::get_compactions.set(r, [&ctx] (std::unique_ptr<http::request> req) {

				        return ctx.db.map_reduce0([](replica::database& db) {

				void set_compaction_manager(http_context& ctx, routes& r, sharded<compaction_manager>& cm) {

				    cm::get_compactions.set(r, [&cm] (std::unique_ptr<http::request> req) {

				        return cm.map_reduce0([](compaction_manager& cm) {

				            std::vector<cm::summary> summaries;

				            const compaction_manager& cm = db.get_compaction_manager();

				            for (const auto& c : cm.get_compactions()) {

				                cm::summary s;

				@@ -100,10 +101,9 @@ void set_compaction_manager(http_context& ctx, routes& r) {

				        return make_ready_future<json::json_return_type>(json_void());

				    });

				    cm::stop_compaction.set(r, [&ctx] (std::unique_ptr<http::request> req) {

				    cm::stop_compaction.set(r, [&cm] (std::unique_ptr<http::request> req) {

				        auto type = req->get_query_param("type");

				        return ctx.db.invoke_on_all([type] (replica::database& db) {

				            auto& cm = db.get_compaction_manager();

				        return cm.invoke_on_all([type] (compaction_manager& cm) {

				            return cm.stop_compaction(type);

				        }).then([] {

				            return make_ready_future<json::json_return_type>(json_void());

				@@ -135,8 +135,8 @@ void set_compaction_manager(http_context& ctx, routes& r) {

				        }, std::plus<int64_t>());

				    });

				    cm::get_completed_tasks.set(r, [&ctx] (std::unique_ptr<http::request> req) {

				        return get_cm_stats(ctx, &compaction_manager::stats::completed_tasks);

				    cm::get_completed_tasks.set(r, [&cm] (std::unique_ptr<http::request> req) {

				        return get_cm_stats(cm, &compaction_manager::stats::completed_tasks);

				    });

				    cm::get_total_compactions_completed.set(r, [] (std::unique_ptr<http::request> req) {

				@@ -153,14 +153,14 @@ void set_compaction_manager(http_context& ctx, routes& r) {

				        return make_ready_future<json::json_return_type>(0);

				    });

				    cm::get_compaction_history.set(r, [&ctx] (std::unique_ptr<http::request> req) {

				        std::function<future<>(output_stream<char>&&)> f = [&ctx] (output_stream<char>&& out) -> future<> {

				    cm::get_compaction_history.set(r, [&cm] (std::unique_ptr<http::request> req) {

				        std::function<future<>(output_stream<char>&&)> f = [&cm] (output_stream<char>&& out) -> future<> {

				            auto s = std::move(out);

				            bool first = true;

				            std::exception_ptr ex;

				            try {

				                co_await s.write("[");

				                co_await ctx.db.local().get_compaction_manager().get_compaction_history([&s, &first](const db::compaction_history_entry& entry) mutable -> future<> {

				                co_await cm.local().get_compaction_history([&s, &first](const db::compaction_history_entry& entry) mutable -> future<> {

				                        cm::history h;

				                        h.id = fmt::to_string(entry.id);

				                        h.ks = std::move(entry.ks);

				@@ -168,7 +168,9 @@ void set_compaction_manager(http_context& ctx, routes& r) {

				                        h.compacted_at = entry.compacted_at;

				                        h.bytes_in = entry.bytes_in;

				                        h.bytes_out =  entry.bytes_out;

				                        for (auto it : entry.rows_merged) {

				                        std::map<int32_t, int64_t> items(entry.rows_merged.begin(), entry.rows_merged.end());

				                        for (auto it : items) {

				                            httpd::compaction_manager_json::row_merged e;

				                            e.key = it.first;

				                            e.value = it.second;

				@@ -201,6 +203,34 @@ void set_compaction_manager(http_context& ctx, routes& r) {

				        return make_ready_future<json::json_return_type>(res);

				    });

				    ss::get_compaction_throughput_mb_per_sec.set(r, [&cm](std::unique_ptr<http::request> req) {

				        int value = cm.local().throughput_mbs();

				        return make_ready_future<json::json_return_type>(value);

				    });

				    ss::set_compaction_throughput_mb_per_sec.set(r, [](std::unique_ptr<http::request> req) {

				        //TBD

				        unimplemented();

				        auto value = req->get_query_param("value");

				        return make_ready_future<json::json_return_type>(json_void());

				    });

				}

				void unset_compaction_manager(http_context& ctx, routes& r) {

				    cm::get_compactions.unset(r);

				    cm::get_pending_tasks_by_table.unset(r);

				    cm::force_user_defined_compaction.unset(r);

				    cm::stop_compaction.unset(r);

				    cm::stop_keyspace_compaction.unset(r);

				    cm::get_pending_tasks.unset(r);

				    cm::get_completed_tasks.unset(r);

				    cm::get_total_compactions_completed.unset(r);

				    cm::get_bytes_compacted.unset(r);

				    cm::get_compaction_history.unset(r);

				    cm::get_compaction_info.unset(r);

				    ss::get_compaction_throughput_mb_per_sec.unset(r);

				    ss::set_compaction_throughput_mb_per_sec.unset(r);

				}

				}

									
										6

api/compaction_manager.hh
									
												View File
												
				@@ -7,13 +7,17 @@

				 */

				#pragma once

				#include <seastar/core/sharded.hh>

				namespace seastar::httpd {

				class routes;

				}

				class compaction_manager;

				namespace api {

				struct http_context;

				void set_compaction_manager(http_context& ctx, seastar::httpd::routes& r);

				void set_compaction_manager(http_context& ctx, seastar::httpd::routes& r, seastar::sharded<compaction_manager>& cm);

				void unset_compaction_manager(http_context& ctx, seastar::httpd::routes& r);

				}

									
										12

api/config.cc
									
												View File
												
				@@ -10,6 +10,7 @@

				#include "api/config.hh"

				#include "api/api-doc/config.json.hh"

				#include "api/api-doc/storage_proxy.json.hh"

				#include "api/api-doc/storage_service.json.hh"

				#include "replica/database.hh"

				#include "db/config.hh"

				#include <sstream>

				@@ -19,6 +20,7 @@

				namespace api {

				using namespace seastar::httpd;

				namespace sp = httpd::storage_proxy_json;

				namespace ss = httpd::storage_service_json;

				template<class T>

				json::json_return_type get_json_return_type(const T& val) {

				@@ -183,6 +185,14 @@ void set_config(std::shared_ptr < api_registry_builder20 > rb, http_context& ctx

				        return make_ready_future<json::json_return_type>(seastar::json::json_void());

				    });

				    ss::get_all_data_file_locations.set(r, [&cfg](const_req req) {

				        return container_to_vec(cfg.data_file_directories());

				    });

				    ss::get_saved_caches_location.set(r, [&cfg](const_req req) {

				        return cfg.saved_caches_directory();

				    });

				}

				void unset_config(http_context& ctx, routes& r) {

				@@ -201,6 +211,8 @@ void unset_config(http_context& ctx, routes& r) {

				    sp::set_range_rpc_timeout.unset(r);

				    sp::get_truncate_rpc_timeout.unset(r);

				    sp::set_truncate_rpc_timeout.unset(r);

				    ss::get_all_data_file_locations.unset(r);

				    ss::get_saved_caches_location.unset(r);

				}

				}

									
										69

api/cql_server_test.cc
									
										Normal file
									
												View File
												
				@@ -0,0 +1,69 @@

				/*

				 * Copyright (C) 2024-present ScyllaDB

				 */

				/*

				 * SPDX-License-Identifier: AGPL-3.0-or-later

				 */

				#ifndef SCYLLA_BUILD_MODE_RELEASE

				#include <seastar/core/coroutine.hh>

				#include "api/api-doc/cql_server_test.json.hh"

				#include "cql_server_test.hh"

				#include "transport/controller.hh"

				#include "transport/server.hh"

				#include "service/qos/qos_common.hh"

				namespace api {

				namespace cst = httpd::cql_server_test_json;

				using namespace json;

				using namespace seastar::httpd;

				struct connection_sl_params : public json::json_base {

				    json::json_element<sstring> _role_name;

				    json::json_element<sstring> _workload_type;

				    json::json_element<sstring> _timeout;

				    connection_sl_params(const sstring& role_name, const sstring& workload_type, const sstring& timeout) {

				        _role_name = role_name;

				        _workload_type = workload_type;

				        _timeout = timeout;

				        register_params();

				    }

				    connection_sl_params(const connection_sl_params& params)

				        : connection_sl_params(params._role_name(), params._workload_type(), params._timeout()) {}

				    void register_params() {

				        add(&_role_name, "role_name");

				        add(&_workload_type, "workload_type");

				        add(&_timeout, "timeout");

				    }    

				};

				void set_cql_server_test(http_context& ctx, seastar::httpd::routes& r, cql_transport::controller& ctl) {

				    cst::connections_params.set(r, [&ctl] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto sl_params = co_await ctl.get_connections_service_level_params();

				        std::vector<connection_sl_params> result;

				        std::ranges::transform(std::move(sl_params), std::back_inserter(result), [] (const cql_transport::connection_service_level_params& params) {

				            auto nanos = std::chrono::duration_cast<std::chrono::nanoseconds>(params.timeout_config.read_timeout).count();

				            return connection_sl_params(

				                    std::move(params.role_name), 

				                    sstring(qos::service_level_options::to_string(params.workload_type)), 

				                    to_string(cql_duration(months_counter{0}, days_counter{0}, nanoseconds_counter{nanos})));

				        });

				        co_return result;

				    });

				}

				void unset_cql_server_test(http_context& ctx, seastar::httpd::routes& r) {

				    cst::connections_params.unset(r);

				}

				}

				#endif

									
										29

api/cql_server_test.hh
									
										Normal file
									
												View File
												
				@@ -0,0 +1,29 @@

				/*

				 * Copyright (C) 2024-present ScyllaDB

				 */

				/*

				 * SPDX-License-Identifier: AGPL-3.0-or-later

				 */

				#ifndef SCYLLA_BUILD_MODE_RELEASE

				#pragma once

				namespace cql_transport {

				class controller;

				}

				namespace seastar::httpd {

				class routes;

				}

				namespace api {

				struct http_context;

				void set_cql_server_test(http_context& ctx, seastar::httpd::routes& r, cql_transport::controller& ctl);

				void unset_cql_server_test(http_context& ctx, seastar::httpd::routes& r);

				}

				#endif

									
										2

api/lsa.cc
									
												View File
												
				@@ -11,7 +11,7 @@

				#include <seastar/http/exception.hh>

				#include "utils/logalloc.hh"

				#include "log.hh"

				#include "utils/log.hh"

				namespace api {

				using namespace seastar::httpd;

									
										48

api/raft.cc
									
												View File
												
				@@ -11,7 +11,8 @@

				#include "api/api-doc/raft.json.hh"

				#include "service/raft/raft_group_registry.hh"

				#include "log.hh"

				#include "service/raft/raft_address_map.hh"

				#include "utils/log.hh"

				using namespace seastar::httpd;

				@@ -102,8 +103,8 @@ void set_raft(http_context&, httpd::routes& r, sharded<service::raft_group_regis

				        if (!req->query_parameters.contains("group_id")) {

				            // Read barrier on group 0 by default

				            co_await raft_gr.invoke_on(0, [timeout] (service::raft_group_registry& raft_gr) {

				                return raft_gr.group0_with_timeouts().read_barrier(nullptr, timeout);

				            co_await raft_gr.invoke_on(0, [timeout] (service::raft_group_registry& raft_gr) -> future<> {

				                co_await raft_gr.group0_with_timeouts().read_barrier(nullptr, timeout);

				            });

				            co_return json_void{};

				        }

				@@ -111,18 +112,52 @@ void set_raft(http_context&, httpd::routes& r, sharded<service::raft_group_regis

				        raft::group_id gid{utils::UUID{req->get_query_param("group_id")}};

				        std::atomic<bool> found_srv{false};

				        co_await raft_gr.invoke_on_all([gid, timeout, &found_srv] (service::raft_group_registry& raft_gr) {

				        co_await raft_gr.invoke_on_all([gid, timeout, &found_srv] (service::raft_group_registry& raft_gr) -> future<> {

				            if (!raft_gr.find_server(gid)) {

				                return make_ready_future<>();

				                co_return;

				            }

				            found_srv = true;

				            return raft_gr.get_server_with_timeouts(gid).read_barrier(nullptr, timeout);

				            co_await raft_gr.get_server_with_timeouts(gid).read_barrier(nullptr, timeout);

				        });

				        if (!found_srv) {

				            throw bad_param_exception{fmt::format("Server for group ID {} not found", gid)};

				        }

				        co_return json_void{};

				    });

				    r::trigger_stepdown.set(r, [&raft_gr] (std::unique_ptr<http::request> req) -> future<json_return_type> {

				        auto timeout = get_request_timeout(*req);

				        auto dur = timeout.value ? *timeout.value - lowres_clock::now() : std::chrono::seconds(60);

				        const auto stepdown_timeout_ticks = dur / service::raft_tick_interval;

				        auto timeout_dur = raft::logical_clock::duration(stepdown_timeout_ticks);

				        if (!req->query_parameters.contains("group_id")) {

				            // Stepdown on group 0 by default

				            co_await raft_gr.invoke_on(0, [timeout_dur] (service::raft_group_registry& raft_gr) {

				                apilog.info("Triggering stepdown for group0");

				                return raft_gr.group0().stepdown(timeout_dur);

				            });

				            co_return json_void{};

				        }

				        raft::group_id gid{utils::UUID{req->get_path_param("group_id")}};

				        std::atomic<bool> found_srv{false};

				        co_await raft_gr.invoke_on_all([gid, timeout_dur, &found_srv] (service::raft_group_registry& raft_gr) -> future<> {

				            auto* srv = raft_gr.find_server(gid);

				            if (!srv) {

				                co_return;

				            }

				            found_srv = true;

				            apilog.info("Triggering stepdown for group {}", gid);

				            co_await srv->stepdown(timeout_dur);

				        });

				        if (!found_srv) {

				            throw std::runtime_error{fmt::format("Server for group ID {} not found", gid)};

				        }

				        co_return json_void{};

				    });

				}

				@@ -131,6 +166,7 @@ void unset_raft(http_context&, httpd::routes& r) {

				    r::trigger_snapshot.unset(r);

				    r::get_leader_host.unset(r);

				    r::read_barrier.unset(r);

				    r::trigger_stepdown.unset(r);

				}

				}

									
										113

api/storage_service.cc
									
												View File
												
				@@ -30,7 +30,6 @@

				#include "service/raft/raft_group0_client.hh"

				#include "service/storage_service.hh"

				#include "service/load_meter.hh"

				#include "db/commitlog/commitlog.hh"

				#include "gms/gossiper.hh"

				#include "db/system_keyspace.hh"

				#include <seastar/http/exception.hh>

				@@ -40,7 +39,7 @@

				#include "repair/row_level.hh"

				#include "locator/snitch_base.hh"

				#include "column_family.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "release.hh"

				#include "compaction/compaction_manager.hh"

				#include "compaction/task_manager_module.hh"

				@@ -54,6 +53,8 @@

				#include "locator/abstract_replication_strategy.hh"

				#include "sstables_loader.hh"

				#include "db/view/view_builder.hh"

				#include "utils/rjson.hh"

				#include "utils/user_provided_param.hh"

				using namespace seastar::httpd;

				using namespace std::chrono_literals;

				@@ -489,10 +490,32 @@ void set_sstables_loader(http_context& ctx, routes& r, sharded<sstables_loader>&

				            return make_ready_future<json::json_return_type>(json_void());

				        });

				    });

				    ss::start_restore.set(r, [&sst_loader] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto endpoint = req->get_query_param("endpoint");

				        auto keyspace = req->get_query_param("keyspace");

				        auto table = req->get_query_param("table");

				        auto bucket = req->get_query_param("bucket");

				        auto prefix = req->get_query_param("prefix");

				        // TODO: the http_server backing the API does not use content streaming

				        // should use it for better performance

				        rjson::value parsed = rjson::parse(req->content);

				        if (!parsed.IsArray()) {

				            throw httpd::bad_param_exception("malformatted sstables in body");

				        }

				        auto sstables = parsed.GetArray() |

				            std::views::transform([] (const auto& s) { return sstring(rjson::to_string_view(s)); }) |

				            std::ranges::to<std::vector>();

				        auto task_id = co_await sst_loader.local().download_new_sstables(keyspace, table, prefix, std::move(sstables), endpoint, bucket);

				        co_return json::json_return_type(fmt::to_string(task_id));

				    });

				}

				void unset_sstables_loader(http_context& ctx, routes& r) {

				    ss::load_new_ss_tables.unset(r);

				    ss::start_restore.unset(r);

				}

				void set_view_builder(http_context& ctx, routes& r, sharded<db::view::view_builder>& vb) {

				@@ -520,10 +543,6 @@ static future<json::json_return_type> describe_ring_as_json_for_table(const shar

				}

				void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_service>& ss, service::raft_group0_client& group0_client) {

				    ss::get_commitlog.set(r, [&ctx](const_req req) {

				        return ctx.db.local().commitlog()->active_config().commit_log_location;

				    });

				    ss::get_token_endpoint.set(r, [&ctx, &ss] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        const auto keyspace_name = req->get_query_param("keyspace");

				        const auto table_name = req->get_query_param("cf");

				@@ -610,14 +629,6 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				        return ss.local().get_schema_version();

				    });

				    ss::get_all_data_file_locations.set(r, [&ctx](const_req req) {

				        return container_to_vec(ctx.db.local().get_config().data_file_directories());

				    });

				    ss::get_saved_caches_location.set(r, [&ctx](const_req req) {

				        return ctx.db.local().get_config().saved_caches_directory();

				    });

				    ss::get_range_to_endpoint_map.set(r, [&ctx, &ss](std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto keyspace = validate_keyspace(ctx, req);

				        auto table = req->get_query_param("cf");

				@@ -706,17 +717,19 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				        auto& db = ctx.db;

				        auto params = req_params({

				            std::pair("flush_memtables", mandatory::no),

				            std::pair("consider_only_existing_data", mandatory::no),

				        });

				        params.process(*req);

				        auto flush = params.get_as<bool>("flush_memtables").value_or(true);

				        apilog.info("force_compaction: flush={}", flush);

				        auto consider_only_existing_data = params.get_as<bool>("consider_only_existing_data").value_or(false);

				        apilog.info("force_compaction: flush={} consider_only_existing_data={}", flush, consider_only_existing_data);

				        auto& compaction_module = db.local().get_compaction_manager().get_task_manager_module();

				        std::optional<flush_mode> fmopt;

				        if (!flush) {

				        if (!flush && !consider_only_existing_data) {

				            fmopt = flush_mode::skip;

				        }

				        auto task = co_await compaction_module.make_and_start_task<global_major_compaction_task_impl>({}, db, fmopt);

				        auto task = co_await compaction_module.make_and_start_task<global_major_compaction_task_impl>({}, db, fmopt, consider_only_existing_data);

				        try {

				            co_await task->done();

				        } catch (...) {

				@@ -733,19 +746,21 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				            std::pair("keyspace", mandatory::yes),

				            std::pair("cf", mandatory::no),

				            std::pair("flush_memtables", mandatory::no),

				            std::pair("consider_only_existing_data", mandatory::no),

				        });

				        params.process(*req);

				        auto keyspace = validate_keyspace(ctx, *params.get("keyspace"));

				        auto table_infos = parse_table_infos(keyspace, ctx, params.get("cf").value_or(""));

				        auto flush = params.get_as<bool>("flush_memtables").value_or(true);

				        apilog.debug("force_keyspace_compaction: keyspace={} tables={}, flush={}", keyspace, table_infos, flush);

				        auto consider_only_existing_data = params.get_as<bool>("consider_only_existing_data").value_or(false);

				        apilog.info("force_keyspace_compaction: keyspace={} tables={}, flush={} consider_only_existing_data={}", keyspace, table_infos, flush, consider_only_existing_data);

				        auto& compaction_module = db.local().get_compaction_manager().get_task_manager_module();

				        std::optional<flush_mode> fmopt;

				        if (!flush) {

				        if (!flush && !consider_only_existing_data) {

				            fmopt = flush_mode::skip;

				        }

				        auto task = co_await compaction_module.make_and_start_task<major_keyspace_compaction_task_impl>({}, std::move(keyspace), tasks::task_id::create_null_id(), db, table_infos, fmopt);

				        auto task = co_await compaction_module.make_and_start_task<major_keyspace_compaction_task_impl>({}, std::move(keyspace), tasks::task_id::create_null_id(), db, table_infos, fmopt, consider_only_existing_data);

				        try {

				            co_await task->done();

				        } catch (...) {

				@@ -884,7 +899,8 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				        auto host_id = validate_host_id(req->get_query_param("host_id"));

				        std::vector<sstring> ignore_nodes_strs = utils::split_comma_separated_list(req->get_query_param("ignore_nodes"));

				        apilog.info("remove_node: host_id={} ignore_nodes={}", host_id, ignore_nodes_strs);

				        auto ignore_nodes = std::list<locator::host_id_or_endpoint>();

				        locator::host_id_or_endpoint_list ignore_nodes;

				        ignore_nodes.reserve(ignore_nodes_strs.size());

				        for (const sstring& n : ignore_nodes_strs) {

				            try {

				                auto hoep = locator::host_id_or_endpoint(n);

				@@ -893,7 +909,7 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				                }

				                ignore_nodes.push_back(std::move(hoep));

				            } catch (...) {

				                throw std::runtime_error(format("Failed to parse ignore_nodes parameter: ignore_nodes={}, node={}: {}", ignore_nodes_strs, n, std::current_exception()));

				                throw std::runtime_error(fmt::format("Failed to parse ignore_nodes parameter: ignore_nodes={}, node={}: {}", ignore_nodes_strs, n, std::current_exception()));

				            }

				        }

				        return ss.local().removenode(host_id, std::move(ignore_nodes)).then([] {

				@@ -1047,18 +1063,6 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				        return make_ready_future<json::json_return_type>(0);

				    });

				    ss::get_compaction_throughput_mb_per_sec.set(r, [&ctx](std::unique_ptr<http::request> req) {

				        int value = ctx.db.local().get_config().compaction_throughput_mb_per_sec();

				        return make_ready_future<json::json_return_type>(value);

				    });

				    ss::set_compaction_throughput_mb_per_sec.set(r, [](std::unique_ptr<http::request> req) {

				        //TBD

				        unimplemented();

				        auto value = req->get_query_param("value");

				        return make_ready_future<json::json_return_type>(json_void());

				    });

				    ss::is_incremental_backups_enabled.set(r, [&ctx](std::unique_ptr<http::request> req) {

				        // If this is issued in parallel with an ongoing change, we may see values not agreeing.

				        // Reissuing is asking for trouble, so we will just return true upon seeing any true value.

				@@ -1096,7 +1100,16 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				    });

				    ss::rebuild.set(r, [&ss](std::unique_ptr<http::request> req) {

				        auto source_dc = req->get_query_param("source_dc");

				        utils::optional_param source_dc;

				        if (auto source_dc_str = req->get_query_param("source_dc"); !source_dc_str.empty()) {

				            source_dc.emplace(std::move(source_dc_str)).set_user_provided();

				        }

				        if (auto force_str = req->get_query_param("force"); !force_str.empty() && service::loosen_constraints(validate_bool(force_str))) {

				            if (!source_dc) {

				                throw bad_param_exception("The `source_dc` option must be provided for using the `force` option");

				            }

				            source_dc.set_force();

				        }

				        apilog.info("rebuild: source_dc={}", source_dc);

				        return ss.local().rebuild(std::move(source_dc)).then([] {

				            return make_ready_future<json::json_return_type>(json_void());

				@@ -1439,12 +1452,7 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				    ss::reload_raft_topology_state.set(r,

				            [&ss, &group0_client] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        co_await ss.invoke_on(0, [&group0_client] (service::storage_service& ss) -> future<> {

				            apilog.info("Waiting for group 0 read/apply mutex before reloading Raft topology state...");

				            auto holder = co_await group0_client.hold_read_apply_mutex();

				            apilog.info("Reloading Raft topology state");

				            // Using topology_transition() instead of topology_state_load(), because the former notifies listeners

				            co_await ss.topology_transition();

				            apilog.info("Reloaded Raft topology state");

				            return ss.reload_raft_topology_state(group0_client);

				        });

				        co_return json_void();

				    });

				@@ -1553,14 +1561,11 @@ void set_storage_service(http_context& ctx, routes& r, sharded<service::storage_

				}

				void unset_storage_service(http_context& ctx, routes& r) {

				    ss::get_commitlog.unset(r);

				    ss::get_token_endpoint.unset(r);

				    ss::toppartitions_generic.unset(r);

				    ss::get_release_version.unset(r);

				    ss::get_scylla_release_version.unset(r);

				    ss::get_schema_version.unset(r);

				    ss::get_all_data_file_locations.unset(r);

				    ss::get_saved_caches_location.unset(r);

				    ss::get_range_to_endpoint_map.unset(r);

				    ss::get_pending_range_to_endpoint_map.unset(r);

				    ss::describe_ring.unset(r);

				@@ -1598,8 +1603,6 @@ void unset_storage_service(http_context& ctx, routes& r) {

				    ss::is_joined.unset(r);

				    ss::set_stream_throughput_mb_per_sec.unset(r);

				    ss::get_stream_throughput_mb_per_sec.unset(r);

				    ss::get_compaction_throughput_mb_per_sec.unset(r);

				    ss::set_compaction_throughput_mb_per_sec.unset(r);

				    ss::is_incremental_backups_enabled.unset(r);

				    ss::set_incremental_backups_enabled.unset(r);

				    ss::rebuild.unset(r);

				@@ -1776,6 +1779,23 @@ void set_snapshot(http_context& ctx, routes& r, sharded<db::snapshot_ctl>& snap_

				        co_return json::json_return_type(static_cast<int>(scrub_status::successful));

				    });

				    ss::start_backup.set(r, [&snap_ctl] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto endpoint = req->get_query_param("endpoint");

				        auto keyspace = req->get_query_param("keyspace");

				        auto table = req->get_query_param("table");

				        auto bucket = req->get_query_param("bucket");

				        auto prefix = req->get_query_param("prefix");

				        auto snapshot_name = req->get_query_param("snapshot");

				        if (snapshot_name.empty()) {

				            // TODO: If missing, snapshot should be taken by scylla, then removed

				            throw httpd::bad_param_exception("The snapshot name must be specified");

				        }

				        auto& ctl = snap_ctl.local();

				        auto task_id = co_await ctl.start_backup(std::move(endpoint), std::move(bucket), std::move(prefix), std::move(keyspace), std::move(table), std::move(snapshot_name));

				        co_return json::json_return_type(fmt::to_string(task_id));

				    });

				    cf::get_true_snapshots_size.set(r, [&snap_ctl] (std::unique_ptr<http::request> req) {

				        auto [ks, cf] = parse_fully_qualified_cf_name(req->get_path_param("name"));

				        return snap_ctl.local().true_snapshots_size(std::move(ks), std::move(cf)).then([] (int64_t res) {

				@@ -1797,6 +1817,7 @@ void unset_snapshot(http_context& ctx, routes& r) {

				    ss::del_snapshot.unset(r);

				    ss::true_snapshots_size.unset(r);

				    ss::scrub.unset(r);

				    ss::start_backup.unset(r);

				    cf::get_true_snapshots_size.unset(r);

				    cf::get_all_true_snapshots_size.unset(r);

				}

									
										1

api/storage_service.hh
									
												View File
												
				@@ -12,6 +12,7 @@

				#include <seastar/json/json_elements.hh>

				#include "api/api_init.hh"

				#include "db/data_listeners.hh"

				#include "compaction/compaction_descriptor.hh"

				namespace cql_transport { class controller; }

				namespace db {

									
										15

api/system.cc
									
												View File
												
				@@ -10,6 +10,7 @@

				#include "api/api-doc/system.json.hh"

				#include "api/api-doc/metrics.json.hh"

				#include "replica/database.hh"

				#include "db/sstables-format-selector.hh"

				#include <rapidjson/document.h>

				#include <seastar/core/reactor.hh>

				@@ -19,7 +20,7 @@

				#include <seastar/util/short_streams.hh>

				#include <seastar/http/short_streams.hh>

				#include "log.hh"

				#include "utils/log.hh"

				extern logging::logger apilog;

				@@ -184,4 +185,16 @@ void set_system(http_context& ctx, routes& r) {

				    }) ;

				}

				void set_format_selector(http_context& ctx, routes& r, db::sstables_format_selector& sel) {

				    hs::get_highest_supported_sstable_version.set(r, [&sel] (std::unique_ptr<request> req) {

				        return smp::submit_to(0, [&sel] {

				            return make_ready_future<json::json_return_type>(seastar::to_sstring(sel.selected_format()));

				        });

				    });

				}

				void unset_format_selector(http_context& ctx, routes& r) {

				    hs::get_highest_supported_sstable_version.unset(r);

				}

				}

									
										5

api/system.hh
									
												View File
												
				@@ -12,9 +12,14 @@ namespace seastar::httpd {

				class routes;

				}

				namespace db { class sstables_format_selector; }

				namespace api {

				struct http_context;

				void set_system(http_context& ctx, seastar::httpd::routes& r);

				void set_format_selector(http_context& ctx, seastar::httpd::routes& r, db::sstables_format_selector& sel);

				void unset_format_selector(http_context& ctx, seastar::httpd::routes& r);

				}

									
										264

api/task_manager.cc
									
												View File
												
				@@ -14,9 +14,10 @@

				#include "api/api.hh"

				#include "api/api-doc/task_manager.json.hh"

				#include "db/system_keyspace.hh"

				#include "tasks/task_handler.hh"

				#include "utils/overloaded_functor.hh"

				#include <utility>

				#include <boost/range/adaptors.hpp>

				namespace api {

				@@ -24,118 +25,87 @@ namespace tm = httpd::task_manager_json;

				using namespace json;

				using namespace seastar::httpd;

				using task_variant = std::variant<tasks::task_manager::foreign_task_ptr, tasks::task_manager::task::task_essentials>;

				inline bool filter_tasks(tasks::task_manager::task_ptr task, std::unordered_map<sstring, sstring>& query_params) {

				    return (!query_params.contains("keyspace") || query_params["keyspace"] == task->get_status().keyspace) &&

				        (!query_params.contains("table") || query_params["table"] == task->get_status().table);

				}

				struct full_task_status {

				    tasks::task_manager::task::status task_status;

				    std::string type;

				    tasks::task_manager::task::progress progress;

				    tasks::task_id parent_id;

				    tasks::is_abortable abortable;

				    std::vector<std::string> children_ids;

				};

				struct task_stats {

				    task_stats(tasks::task_manager::task_ptr task)

				        : task_id(task->id().to_sstring())

				        , state(task->get_status().state)

				        , type(task->type())

				        , scope(task->get_status().scope)

				        , keyspace(task->get_status().keyspace)

				        , table(task->get_status().table)

				        , entity(task->get_status().entity)

				        , sequence_number(task->get_status().sequence_number)

				    { }

				    sstring task_id;

				    tasks::task_manager::task_state state;

				    std::string type;

				    std::string scope;

				    std::string keyspace;

				    std::string table;

				    std::string entity;

				    uint64_t sequence_number;

				};

				tm::task_status make_status(full_task_status status) {

				    auto start_time = db_clock::to_time_t(status.task_status.start_time);

				    auto end_time = db_clock::to_time_t(status.task_status.end_time);

				tm::task_status make_status(tasks::task_status status) {

				    auto start_time = db_clock::to_time_t(status.start_time);

				    auto end_time = db_clock::to_time_t(status.end_time);

				    ::tm st, et;

				    ::gmtime_r(&end_time, &et);

				    ::gmtime_r(&start_time, &st);

				    std::vector<tm::task_identity> tis{status.children.size()};

				    std::ranges::transform(status.children, tis.begin(), [] (const auto& child) {

				        tm::task_identity ident;

				        ident.task_id = child.task_id.to_sstring();

				        ident.node = fmt::format("{}", child.node);

				        return ident;

				    });

				    tm::task_status res{};

				    res.id = status.task_status.id.to_sstring();

				    res.id = status.task_id.to_sstring();

				    res.type = status.type;

				    res.scope = status.task_status.scope;

				    res.state = status.task_status.state;

				    res.is_abortable = bool(status.abortable);

				    res.kind = status.kind;

				    res.scope = status.scope;

				    res.state = status.state;

				    res.is_abortable = bool(status.is_abortable);

				    res.start_time = st;

				    res.end_time = et;

				    res.error = status.task_status.error;

				    res.parent_id = status.parent_id.to_sstring();

				    res.sequence_number = status.task_status.sequence_number;

				    res.shard = status.task_status.shard;

				    res.keyspace = status.task_status.keyspace;

				    res.table = status.task_status.table;

				    res.entity = status.task_status.entity;

				    res.progress_units = status.task_status.progress_units;

				    res.error = status.error;

				    res.parent_id = status.parent_id ? status.parent_id.to_sstring() : "none";

				    res.sequence_number = status.sequence_number;

				    res.shard = status.shard;

				    res.keyspace = status.keyspace;

				    res.table = status.table;

				    res.entity = status.entity;

				    res.progress_units = status.progress_units;

				    res.progress_total = status.progress.total;

				    res.progress_completed = status.progress.completed;

				    res.children_ids = std::move(status.children_ids);

				    res.children_ids = std::move(tis);

				    return res;

				}

				future<full_task_status> retrieve_status(const tasks::task_manager::foreign_task_ptr& task) {

				    if (task.get() == nullptr) {

				        co_return coroutine::return_exception(httpd::bad_param_exception("Task not found"));

				    }

				    auto progress = co_await task->get_progress();

				    full_task_status s;

				    s.task_status = task->get_status();

				    s.type = task->type();

				    s.parent_id = task->get_parent_id();

				    s.abortable = task->is_abortable();

				    s.progress.completed = progress.completed;

				    s.progress.total = progress.total;

				    std::vector<std::string> ct = co_await task->get_children().map_each_task<std::string>([] (const tasks::task_manager::foreign_task_ptr& child) {

				        return child->id().to_sstring();

				    }, [] (const tasks::task_manager::task::task_essentials& child) {

				        return child.task_status.id.to_sstring();

				    });

				    s.children_ids = std::move(ct);

				    co_return s;

				};

				tm::task_stats make_stats(tasks::task_stats stats) {

				    tm::task_stats res{};

				    res.task_id = stats.task_id.to_sstring();

				    res.type = stats.type;

				    res.kind = stats.kind;

				    res.scope = stats.scope;

				    res.state = stats.state;

				    res.sequence_number = stats.sequence_number;

				    res.keyspace = stats.keyspace;

				    res.table = stats.table;

				    res.entity = stats.entity;

				    return res;

				}

				void set_task_manager(http_context& ctx, routes& r, sharded<tasks::task_manager>& tm, db::config& cfg) {

				    tm::get_modules.set(r, [&tm] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        std::vector<std::string> v = boost::copy_range<std::vector<std::string>>(tm.local().get_modules() | boost::adaptors::map_keys);

				        std::vector<std::string> v = tm.local().get_modules() | std::views::keys | std::ranges::to<std::vector>();

				        co_return v;

				    });

				    tm::get_tasks.set(r, [&tm] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        using chunked_stats = utils::chunked_vector<task_stats>;

				        using chunked_stats = utils::chunked_vector<tasks::task_stats>;

				        auto internal = tasks::is_internal{req_param<bool>(*req, "internal", false)};

				        std::vector<chunked_stats> res = co_await tm.map([&req, internal] (tasks::task_manager& tm) {

				            chunked_stats local_res;

				            tasks::task_manager::module_ptr module;

				            std::optional<std::string> keyspace = std::nullopt;

				            std::optional<std::string> table = std::nullopt;

				            try {

				                module = tm.find_module(req->get_path_param("module"));

				            } catch (...) {

				                throw bad_param_exception(fmt::format("{}", std::current_exception()));

				            }

				            const auto& filtered_tasks = module->get_tasks() | boost::adaptors::filtered([&params = req->query_parameters, internal] (const auto& task) {

				                return (internal || !task.second->is_internal()) && filter_tasks(task.second, params);

				            });

				            for (auto& [task_id, task] : filtered_tasks) {

				                local_res.push_back(task_stats{task});

				            if (auto it = req->query_parameters.find("keyspace"); it != req->query_parameters.end()) {

				                keyspace = it->second;

				            }

				            return local_res;

				            if (auto it = req->query_parameters.find("table"); it != req->query_parameters.end()) {

				                table = it->second;

				            }

				            return module->get_stats(internal, [keyspace = std::move(keyspace), table = std::move(table)] (std::string& ks, std::string& t) {

				                return (!keyspace || keyspace == ks) && (!table || table == t);

				            });

				        });

				        std::function<future<>(output_stream<char>&&)> f = [r = std::move(res)] (output_stream<char>&& os) -> future<> {

				@@ -148,8 +118,7 @@ void set_task_manager(http_context& ctx, routes& r, sharded<tasks::task_manager>

				                for (auto& v: res) {

				                    for (auto& stats: v) {

				                        co_await s.write(std::exchange(delim, ", "));

				                        tm::task_stats ts;

				                        ts = stats;

				                        tm::task_stats ts = make_stats(stats);

				                        co_await formatter::write(s, ts);

				                    }

				                }

				@@ -168,121 +137,70 @@ void set_task_manager(http_context& ctx, routes& r, sharded<tasks::task_manager>

				    tm::get_task_status.set(r, [&tm] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto id = tasks::task_id{utils::UUID{req->get_path_param("task_id")}};

				        tasks::task_manager::foreign_task_ptr task;

				        tasks::task_status status;

				        try {

				            task = co_await tasks::task_manager::invoke_on_task(tm, id, std::function([] (tasks::task_manager::task_ptr task) -> future<tasks::task_manager::foreign_task_ptr> {

				                if (task->is_complete()) {

				                    task->unregister_task();

				                }

				                co_return std::move(task);

				            }));

				            auto task = tasks::task_handler{tm.local(), id};

				            status = co_await task.get_status();

				        } catch (tasks::task_manager::task_not_found& e) {

				            throw bad_param_exception(e.what());

				        }

				        auto s = co_await retrieve_status(task);

				        co_return make_status(s);

				        co_return make_status(status);

				    });

				    tm::abort_task.set(r, [&tm] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto id = tasks::task_id{utils::UUID{req->get_path_param("task_id")}};

				        try {

				            co_await tasks::task_manager::invoke_on_task(tm, id, [] (tasks::task_manager::task_ptr task) -> future<> {

				                if (!task->is_abortable()) {

				                    co_await coroutine::return_exception(std::runtime_error("Requested task cannot be aborted"));

				                }

				                task->abort();

				            });

				            auto task = tasks::task_handler{tm.local(), id};

				            co_await task.abort();

				        } catch (tasks::task_manager::task_not_found& e) {

				            throw bad_param_exception(e.what());

				        } catch (tasks::task_not_abortable& e) {

				            throw httpd::base_exception{e.what(), http::reply::status_type::forbidden};

				        }

				        co_return json_void();

				    });

				    tm::wait_task.set(r, [&tm] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto id = tasks::task_id{utils::UUID{req->get_path_param("task_id")}};

				        tasks::task_manager::foreign_task_ptr task;

				        tasks::task_status status;

				        std::optional<std::chrono::seconds> timeout = std::nullopt;

				        if (auto it = req->query_parameters.find("timeout"); it != req->query_parameters.end()) {

				            timeout = std::chrono::seconds(boost::lexical_cast<uint32_t>(it->second));

				        }

				        try {

				            task = co_await tasks::task_manager::invoke_on_task(tm, id, std::function([] (tasks::task_manager::task_ptr task) {

				                return task->done().then_wrapped([task] (auto f) {

				                    // done() is called only because we want the task to be complete before getting its status.

				                    // The future should be ignored here as the result does not matter.

				                    f.ignore_ready_future();

				                    return make_foreign(task);

				                });

				            }));

				            auto task = tasks::task_handler{tm.local(), id};

				            status = co_await task.wait_for_task(timeout);

				        } catch (tasks::task_manager::task_not_found& e) {

				            throw bad_param_exception(e.what());

				        } catch (timed_out_error& e) {

				            throw httpd::base_exception{e.what(), http::reply::status_type::request_timeout};

				        }

				        auto s = co_await retrieve_status(task);

				        co_return make_status(s);

				        co_return make_status(status);

				    });

				    tm::get_task_status_recursively.set(r, [&_tm = tm] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto& tm = _tm;

				        auto id = tasks::task_id{utils::UUID{req->get_path_param("task_id")}};

				        std::queue<task_variant> q;

				        utils::chunked_vector<full_task_status> res;

				        tasks::task_manager::foreign_task_ptr task;

				        try {

				            // Get requested task.

				            task = co_await tasks::task_manager::invoke_on_task(tm, id, std::function([] (tasks::task_manager::task_ptr task) -> future<tasks::task_manager::foreign_task_ptr> {

				                if (task->is_complete()) {

				                    task->unregister_task();

				            auto task = tasks::task_handler{tm.local(), id};

				            auto res = co_await task.get_status_recursively(true);

				            std::function<future<>(output_stream<char>&&)> f = [r = std::move(res)] (output_stream<char>&& os) -> future<> {

				                auto s = std::move(os);

				                auto res = std::move(r);

				                co_await s.write("[");

				                std::string delim = "";

				                for (auto& status: res) {

				                    co_await s.write(std::exchange(delim, ", "));

				                    co_await formatter::write(s, make_status(status));

				                }

				                co_return task;

				            }));

				                co_await s.write("]");

				                co_await s.close();

				            };

				            co_return f;

				        } catch (tasks::task_manager::task_not_found& e) {

				            throw bad_param_exception(e.what());

				        }

				        // Push children's statuses in BFS order.

				        q.push(co_await task.copy());   // Task cannot be moved since we need it to be alive during whole loop execution.

				        while (!q.empty()) {

				            auto& current = q.front();

				            co_await std::visit(overloaded_functor {

				                [&] (const tasks::task_manager::foreign_task_ptr& task) -> future<> {

				                    res.push_back(co_await retrieve_status(task));

				                    co_await task->get_children().for_each_task([&q] (const tasks::task_manager::foreign_task_ptr& child) -> future<> {

				                        q.push(co_await child.copy());

				                    }, [&] (const tasks::task_manager::task::task_essentials& child) {

				                        q.push(child);

				                        return make_ready_future();

				                    });

				                },

				                [&] (const tasks::task_manager::task::task_essentials& task) -> future<> {

				                    res.push_back(full_task_status{

				                        .task_status = task.task_status,

				                        .type = task.type,

				                        .progress = task.task_progress,

				                        .parent_id = task.parent_id,

				                        .abortable = task.abortable,

				                        .children_ids = boost::copy_range<std::vector<std::string>>(task.failed_children | boost::adaptors::transformed([] (auto& child) {

				                            return child.task_status.id.to_sstring();

				                        }))

				                    });

				                    for (auto& child: task.failed_children) {

				                        q.push(child);

				                    }

				                    return make_ready_future();

				                }

				            }, current);

				            q.pop();

				        }

				        std::function<future<>(output_stream<char>&&)> f = [r = std::move(res)] (output_stream<char>&& os) -> future<> {

				            auto s = std::move(os);

				            auto res = std::move(r);

				            co_await s.write("[");

				            std::string delim = "";

				            for (auto& status: res) {

				                co_await s.write(std::exchange(delim, ", "));

				                co_await formatter::write(s, make_status(status));

				            }

				            co_await s.write("]");

				            co_await s.close();

				        };

				        co_return f;

				    });

				    tm::get_and_update_ttl.set(r, [&cfg] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				@@ -294,6 +212,11 @@ void set_task_manager(http_context& ctx, routes& r, sharded<tasks::task_manager>

				        }

				        co_return json::json_return_type(ttl);

				    });

				    tm::get_ttl.set(r, [&cfg] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        uint32_t ttl = cfg.task_ttl_seconds();

				        co_return json::json_return_type(ttl);

				    });

				}

				void unset_task_manager(http_context& ctx, routes& r) {

				@@ -304,6 +227,7 @@ void unset_task_manager(http_context& ctx, routes& r) {

				    tm::wait_task.unset(r);

				    tm::get_task_status_recursively.unset(r);

				    tm::get_and_update_ttl.unset(r);

				    tm::get_ttl.unset(r);

				}

				}

									
										39

api/task_manager_test.cc
									
												View File
												
				@@ -13,6 +13,7 @@

				#include "task_manager_test.hh"

				#include "api/api-doc/task_manager_test.json.hh"

				#include "tasks/test_module.hh"

				#include "utils/overloaded_functor.hh"

				namespace api {

				@@ -61,8 +62,8 @@ void set_task_manager_test(http_context& ctx, routes& r, sharded<tasks::task_man

				        auto module = tms.local().find_module("test");

				        id = co_await module->make_task<tasks::test_task_impl>(shard, id, keyspace, table, entity, data);

				        co_await tms.invoke_on(shard, [id] (tasks::task_manager& tm) {

				            auto it = tm.get_all_tasks().find(id);

				            if (it != tm.get_all_tasks().end()) {

				            auto it = tm.get_local_tasks().find(id);

				            if (it != tm.get_local_tasks().end()) {

				                it->second->start();

				            }

				        });

				@@ -72,9 +73,16 @@ void set_task_manager_test(http_context& ctx, routes& r, sharded<tasks::task_man

				    tmt::unregister_test_task.set(r, [&tm] (std::unique_ptr<http::request> req) -> future<json::json_return_type> {

				        auto id = tasks::task_id{utils::UUID{req->query_parameters["task_id"]}};

				        try {

				            co_await tasks::task_manager::invoke_on_task(tm, id, [] (tasks::task_manager::task_ptr task) -> future<> {

				                tasks::test_task test_task{task};

				                co_await test_task.unregister_task();

				            co_await tasks::task_manager::invoke_on_task(tm, id, [] (tasks::task_manager::task_variant task_v) -> future<> {

				                return std::visit(overloaded_functor{

				                    [] (tasks::task_manager::task_ptr task) -> future<> {

				                        tasks::test_task test_task{task};

				                        co_await test_task.unregister_task();

				                    },

				                    [] (tasks::task_manager::virtual_task_ptr task) {

				                        return make_ready_future();

				                    }

				                }, task_v);

				            });

				        } catch (tasks::task_manager::task_not_found& e) {

				            throw bad_param_exception(e.what());

				@@ -89,13 +97,20 @@ void set_task_manager_test(http_context& ctx, routes& r, sharded<tasks::task_man

				        std::string error = fail ? it->second : "";

				        try {

				            co_await tasks::task_manager::invoke_on_task(tm, id, [fail, error = std::move(error)] (tasks::task_manager::task_ptr task) -> future<> {

				                tasks::test_task test_task{task};

				                if (fail) {

				                    co_await test_task.finish_failed(std::make_exception_ptr(std::runtime_error(error)));

				                } else {

				                    co_await test_task.finish();

				                }

				            co_await tasks::task_manager::invoke_on_task(tm, id, [fail, error = std::move(error)] (tasks::task_manager::task_variant task_v) -> future<> {

				                return std::visit(overloaded_functor{

				                    [fail, error = std::move(error)] (tasks::task_manager::task_ptr task) -> future<> {

				                        tasks::test_task test_task{task};

				                        if (fail) {

				                            co_await test_task.finish_failed(std::make_exception_ptr(std::runtime_error(error)));

				                        } else {

				                            co_await test_task.finish();

				                        }

				                    },

				                    [] (tasks::task_manager::virtual_task_ptr task) {

				                        return make_ready_future();

				                    }

				                }, task_v);

				            });

				        } catch (tasks::task_manager::task_not_found& e) {

				            throw bad_param_exception(e.what());

									
										1

api/tasks.cc
									
												View File
												
				@@ -16,6 +16,7 @@

				#include "compaction/task_manager_module.hh"

				#include "service/storage_service.hh"

				#include "tasks/task_manager.hh"

				#include "replica/database.hh"

				using namespace seastar::httpd;

									
										5

api/token_metadata.cc
									
												View File
												
				@@ -21,6 +21,9 @@ using namespace json;

				void set_token_metadata(http_context& ctx, routes& r, sharded<locator::shared_token_metadata>& tm) {

				    ss::local_hostid.set(r, [&tm](std::unique_ptr<http::request> req) {

				        auto id = tm.local().get()->get_my_id();

				        if (!bool(id)) {

				            throw not_found_exception("local host ID is not yet set");

				        }

				        return make_ready_future<json::json_return_type>(id.to_sstring());

				    });

				@@ -68,7 +71,7 @@ void set_token_metadata(http_context& ctx, routes& r, sharded<locator::shared_to

				    ss::get_host_id_map.set(r, [&tm](const_req req) {

				        std::vector<ss::mapper> res;

				        return map_to_key_value(tm.local().get()->get_endpoint_to_host_id_map_for_reading(), res);

				        return map_to_key_value(tm.local().get()->get_endpoint_to_host_id_map(), res);

				    });

				    static auto host_or_broadcast = [&tm](const_req req) {

									
										57

auth/authentication_options.hh
									
												View File
												
				@@ -12,6 +12,7 @@

				#include <stdexcept>

				#include <unordered_map>

				#include <unordered_set>

				#include <variant>

				#include <seastar/core/print.hh>

				#include <seastar/core/sstring.hh>

				@@ -22,29 +23,10 @@ namespace auth {

				enum class authentication_option {

				    password,

				    hashed_password,

				    options

				};

				using authentication_option_set = std::unordered_set<authentication_option>;

				using custom_options = std::unordered_map<sstring, sstring>;

				struct authentication_options final {

				    std::optional<sstring> password;

				    std::optional<custom_options> options;

				};

				inline bool any_authentication_options(const authentication_options& aos) noexcept {

				    return aos.password || aos.options;

				}

				class unsupported_authentication_option : public std::invalid_argument {

				public:

				    explicit unsupported_authentication_option(authentication_option k)

				            : std::invalid_argument(format("The {} option is not supported.", k)) {

				    }

				};

				}

				template <>

				@@ -55,9 +37,44 @@ struct fmt::formatter<auth::authentication_option> : fmt::formatter<string_view>

				        switch (a) {

				        case password:

				            return formatter<string_view>::format("PASSWORD", ctx);

				        case hashed_password:

				            return formatter<string_view>::format("HASHED PASSWORD", ctx);

				        case options:

				            return formatter<string_view>::format("OPTIONS", ctx);

				        }

				        std::abort();

				    }

				};

				namespace auth {

				using authentication_option_set = std::unordered_set<authentication_option>;

				using custom_options = std::unordered_map<sstring, sstring>;

				struct password_option {

				    sstring password;

				};

				/// Used exclusively for restoring roles.

				struct hashed_password_option {

				    sstring hashed_password;

				};

				struct authentication_options final {

				    std::optional<std::variant<password_option, hashed_password_option>> credentials;

				    std::optional<custom_options> options;

				};

				inline bool any_authentication_options(const authentication_options& aos) noexcept {

				    return aos.options || aos.credentials;

				}

				class unsupported_authentication_option : public std::invalid_argument {

				public:

				    explicit unsupported_authentication_option(authentication_option k)

				            : std::invalid_argument(format("The {} option is not supported.", k)) {

				    }

				};

				}

									
										18

auth/authenticator.hh
									
												View File
												
				@@ -43,6 +43,11 @@ struct certificate_info {

				using session_dn_func = std::function<future<std::optional<certificate_info>>()>;

				class unsupported_authentication_operation : public std::invalid_argument {

				public:

				    using std::invalid_argument::invalid_argument;

				};

				///

				/// Abstract client for authenticating role identity.

				///

				@@ -129,6 +134,19 @@ public:

				    ///

				    virtual future<custom_options> query_custom_options(std::string_view role_name) const = 0;

				    virtual bool uses_password_hashes() const {

				        return false;

				    }

				    ///

				    /// Query the password hash corresponding to a given role.

				    ///

				    /// If the authenticator doesn't use password hashes, throws an `unsupported_authentication_operation` exception.

				    ///

				    virtual future<std::optional<sstring>> get_password_hash(std::string_view role_name) const {

				        return make_exception_future<std::optional<sstring>>(unsupported_authentication_operation("get_password_hash is not implemented"));

				    }

				    ///

				    /// System resources used internally as part of the implementation. These are made inaccessible to users.

				    ///

									
										6

auth/certificate_authenticator.cc
									
												View File
												
				@@ -9,7 +9,7 @@

				#include "auth/certificate_authenticator.hh"

				#include <regex>

				#include <boost/regex.hpp>

				#include <fmt/ranges.h>

				#include "utils/class_registrator.hh"

				@@ -76,7 +76,7 @@ auth::certificate_authenticator::certificate_authenticator(cql3::query_processor

				                    continue;

				                } catch (std::out_of_range&) {

				                    // just fallthrough

				                } catch (std::regex_error&) {

				                } catch (boost::regex_error&) {

				                    std::throw_with_nested(std::invalid_argument(fmt::format("Invalid query expression: {}", map.at(cfg_query_attr))));

				                }

				            }

				@@ -149,7 +149,7 @@ future<std::optional<auth::authenticated_user>> auth::certificate_authenticator:

				            co_return username;

				        }

				    }

				    throw exceptions::authentication_exception(format("Subject '{}'/'{}' does not match any query expression", subject, altname));

				    throw exceptions::authentication_exception(seastar::format("Subject '{}'/'{}' does not match any query expression", subject, altname));

				}

									
										1

auth/certificate_authenticator.hh
									
												View File
												
				@@ -10,6 +10,7 @@

				#pragma once

				#include "auth/authenticator.hh"

				#include <boost/regex_fwd.hpp>  // IWYU pragma: keep

				namespace cql3 {

									
										14

auth/common.cc
									
												View File
												
				@@ -16,6 +16,7 @@

				#include "mutation/canonical_mutation.hh"

				#include "schema/schema_fwd.hh"

				#include "timestamp.hh"

				#include "utils/assert.hh"

				#include "utils/exponential_backoff_retry.hh"

				#include "cql3/query_processor.hh"

				#include "cql3/statements/create_table_statement.hh"

				@@ -24,6 +25,7 @@

				#include "service/raft/group0_state_machine.hh"

				#include "timeout_config.hh"

				#include "utils/error_injection.hh"

				#include "db/system_keyspace.hh"

				namespace auth {

				@@ -39,7 +41,7 @@ constinit const std::string_view AUTH_PACKAGE_NAME("org.apache.cassandra.auth.")

				static logging::logger auth_log("auth");

				bool legacy_mode(cql3::query_processor& qp) {

				    return qp.auth_version < db::system_keyspace::auth_version_t::v2;

				    return qp.auth_version < db::auth_version_t::v2;

				}

				std::string_view get_auth_ks_name(cql3::query_processor& qp) {

				@@ -68,10 +70,10 @@ static future<> create_legacy_metadata_table_if_missing_impl(

				        cql3::query_processor& qp,

				        std::string_view cql,

				        ::service::migration_manager& mm) {

				    assert(this_shard_id() == 0); // once_among_shards makes sure a function is executed on shard 0 only

				    SCYLLA_ASSERT(this_shard_id() == 0); // once_among_shards makes sure a function is executed on shard 0 only

				    auto db = qp.db();

				    auto parsed_statement = cql3::query_processor::parse_statement(cql);

				    auto parsed_statement = cql3::query_processor::parse_statement(cql, cql3::dialect{});

				    auto& parsed_cf_statement = static_cast<cql3::statements::raw::cf_statement&>(*parsed_statement);

				    parsed_cf_statement.prepare_keyspace(meta::legacy::AUTH_KS);

				@@ -121,7 +123,7 @@ static future<> announce_mutations_with_guard(

				        ::service::raft_group0_client& group0_client,

				        std::vector<canonical_mutation> muts,

				        ::service::group0_guard group0_guard,

				        seastar::abort_source* as,

				        seastar::abort_source& as,

				        std::optional<::service::raft_timeout> timeout) {

				    auto group0_cmd = group0_client.prepare_command(

				        ::service::write_mutations{

				@@ -137,7 +139,7 @@ future<> announce_mutations_with_batching(

				        ::service::raft_group0_client& group0_client,

				        start_operation_func_t start_operation_func,

				        std::function<::service::mutations_generator(api::timestamp_type t)> gen,

				        seastar::abort_source* as,

				        seastar::abort_source& as,

				        std::optional<::service::raft_timeout> timeout) {

				    // account for command's overhead, it's better to use smaller threshold than constantly bounce off the limit

				    size_t memory_threshold = group0_client.max_command_size() * 0.75;

				@@ -188,7 +190,7 @@ future<> announce_mutations(

				        ::service::raft_group0_client& group0_client,

				        const sstring query_string,

				        std::vector<data_value_or_unset> values,

				        seastar::abort_source* as,

				        seastar::abort_source& as,

				        std::optional<::service::raft_timeout> timeout) {

				    auto group0_guard = co_await group0_client.start_operation(as, timeout);

				    auto timestamp = group0_guard.write_timestamp();

									
										6

auth/common.hh
									
												View File
												
				@@ -80,7 +80,7 @@ future<> create_legacy_metadata_table_if_missing(

				// Execute update query via group0 mechanism, mutations will be applied on all nodes.

				// Use this function when need to perform read before write on a single guard or if

				// you have more than one mutation and potentially exceed single command size limit.

				using start_operation_func_t = std::function<future<::service::group0_guard>(abort_source*)>;

				using start_operation_func_t = std::function<future<::service::group0_guard>(abort_source&)>;

				future<> announce_mutations_with_batching(

				        ::service::raft_group0_client& group0_client,

				        // since we can operate also in topology coordinator context where we need stronger

				@@ -88,7 +88,7 @@ future<> announce_mutations_with_batching(

				        // function here

				        start_operation_func_t start_operation_func,

				        std::function<::service::mutations_generator(api::timestamp_type t)> gen,

				        seastar::abort_source* as,

				        seastar::abort_source& as,

				        std::optional<::service::raft_timeout> timeout);

				// Execute update query via group0 mechanism, mutations will be applied on all nodes.

				@@ -97,7 +97,7 @@ future<> announce_mutations(

				        ::service::raft_group0_client& group0_client,

				        const sstring query_string,

				        std::vector<data_value_or_unset> values,

				        seastar::abort_source* as,

				        seastar::abort_source& as,

				        std::optional<::service::raft_timeout> timeout);

				// Appends mutations to a collector, they will be applied later on all nodes via group0 mechanism.

									
										28

auth/default_authorizer.cc
									
												View File
												
				@@ -27,7 +27,7 @@ extern "C" {

				#include "cql3/query_processor.hh"

				#include "cql3/untyped_result_set.hh"

				#include "exceptions/exceptions.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "utils/class_registrator.hh"

				namespace auth {

				@@ -67,7 +67,7 @@ bool default_authorizer::legacy_metadata_exists() const {

				}

				future<bool> default_authorizer::legacy_any_granted() const {

				    static const sstring query = format("SELECT * FROM {}.{} LIMIT 1", meta::legacy::AUTH_KS, PERMISSIONS_CF);

				    static const sstring query = seastar::format("SELECT * FROM {}.{} LIMIT 1", meta::legacy::AUTH_KS, PERMISSIONS_CF);

				    return _qp.execute_internal(

				            query,

				@@ -80,7 +80,7 @@ future<bool> default_authorizer::legacy_any_granted() const {

				future<> default_authorizer::migrate_legacy_metadata() {

				    alogger.info("Starting migration of legacy permissions metadata.");

				    static const sstring query = format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, legacy_table_name);

				    static const sstring query = seastar::format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, legacy_table_name);

				    return _qp.execute_internal(

				            query,

				@@ -163,7 +163,7 @@ default_authorizer::authorize(const role_or_anonymous& maybe_role, const resourc

				        co_return permissions::NONE;

				    }

				    const sstring query = format("SELECT {} FROM {}.{} WHERE {} = ? AND {} = ?",

				    const sstring query = seastar::format("SELECT {} FROM {}.{} WHERE {} = ? AND {} = ?",

				            PERMISSIONS_NAME,

				            get_auth_ks_name(_qp),

				            PERMISSIONS_CF,

				@@ -188,7 +188,7 @@ default_authorizer::modify(

				        const resource& resource,

				        std::string_view op,

				        ::service::group0_batch& mc) {

				    const sstring query = format("UPDATE {}.{} SET {} = {} {} ? WHERE {} = ? AND {} = ?",

				    const sstring query = seastar::format("UPDATE {}.{} SET {} = {} {} ? WHERE {} = ? AND {} = ?",

				            get_auth_ks_name(_qp),

				            PERMISSIONS_CF,

				            PERMISSIONS_NAME,

				@@ -218,7 +218,7 @@ future<> default_authorizer::revoke(std::string_view role_name, permission_set s

				}

				future<std::vector<permission_details>> default_authorizer::list_all() const {

				    const sstring query = format("SELECT {}, {}, {} FROM {}.{}",

				    const sstring query = seastar::format("SELECT {}, {}, {} FROM {}.{}",

				            ROLE_NAME,

				            RESOURCE_NAME,

				            PERMISSIONS_NAME,

				@@ -246,7 +246,7 @@ future<std::vector<permission_details>> default_authorizer::list_all() const {

				future<> default_authorizer::revoke_all(std::string_view role_name, ::service::group0_batch& mc) {

				    try {

				        const sstring query = format("DELETE FROM {}.{} WHERE {} = ?",

				        const sstring query = seastar::format("DELETE FROM {}.{} WHERE {} = ?",

				                get_auth_ks_name(_qp),

				                PERMISSIONS_CF,

				                ROLE_NAME);

				@@ -266,7 +266,7 @@ future<> default_authorizer::revoke_all(std::string_view role_name, ::service::g

				}

				future<> default_authorizer::revoke_all_legacy(const resource& resource) {

				    static const sstring query = format("SELECT {} FROM {}.{} WHERE {} = ? ALLOW FILTERING",

				    static const sstring query = seastar::format("SELECT {} FROM {}.{} WHERE {} = ? ALLOW FILTERING",

				            ROLE_NAME,

				            get_auth_ks_name(_qp),

				            PERMISSIONS_CF,

				@@ -283,7 +283,7 @@ future<> default_authorizer::revoke_all_legacy(const resource& resource) {

				                    res->begin(),

				                    res->end(),

				                    [this, res, resource](const cql3::untyped_result_set::row& r) {

				                static const sstring query = format("DELETE FROM {}.{} WHERE {} = ? AND {} = ?",

				                static const sstring query = seastar::format("DELETE FROM {}.{} WHERE {} = ? AND {} = ?",

				                        get_auth_ks_name(_qp),

				                        PERMISSIONS_CF,

				                        ROLE_NAME,

				@@ -323,7 +323,7 @@ future<> default_authorizer::revoke_all(const resource& resource, ::service::gro

				    auto name = resource.name();

				    auto gen = [this, name] (api::timestamp_type t) -> ::service::mutations_generator {

				        const sstring query = format("SELECT {} FROM {}.{} WHERE {} = ? ALLOW FILTERING",

				        const sstring query = seastar::format("SELECT {} FROM {}.{} WHERE {} = ? ALLOW FILTERING",

				                ROLE_NAME,

				                get_auth_ks_name(_qp),

				                PERMISSIONS_CF,

				@@ -334,7 +334,7 @@ future<> default_authorizer::revoke_all(const resource& resource, ::service::gro

				                {name},

				                cql3::query_processor::cache_internal::no);

				        for (const auto& r : *res) {

				            const sstring query = format("DELETE FROM {}.{} WHERE {} = ? AND {} = ?",

				            const sstring query = seastar::format("DELETE FROM {}.{} WHERE {} = ? AND {} = ?",

				                    get_auth_ks_name(_qp),

				                    PERMISSIONS_CF,

				                    ROLE_NAME,

				@@ -346,7 +346,7 @@ future<> default_authorizer::revoke_all(const resource& resource, ::service::gro

				                    {r.get_as<sstring>(ROLE_NAME), name});

				            if (muts.size() != 1) {

				                on_internal_error(alogger,

				                    format("expecting single delete mutation, got {}", muts.size()));

				                    seastar::format("expecting single delete mutation, got {}", muts.size()));

				            }

				            co_yield std::move(muts[0]);

				        }

				@@ -357,7 +357,7 @@ future<> default_authorizer::revoke_all(const resource& resource, ::service::gro

				void default_authorizer::revoke_all_keyspace_resources(const resource& ks_resource, ::service::group0_batch& mc) {

				    auto ks_name = ks_resource.name();

				    auto gen = [this, ks_name] (api::timestamp_type t) -> ::service::mutations_generator {

				        const sstring query = format("SELECT {}, {} FROM {}.{}",

				        const sstring query = seastar::format("SELECT {}, {} FROM {}.{}",

				                ROLE_NAME,

				                RESOURCE_NAME,

				                get_auth_ks_name(_qp),

				@@ -374,7 +374,7 @@ void default_authorizer::revoke_all_keyspace_resources(const resource& ks_resour

				                // r doesn't represent resource related to ks_resource

				                continue;

				            }

				            const sstring query = format("DELETE FROM {}.{} WHERE {} = ? AND {} = ?",

				            const sstring query = seastar::format("DELETE FROM {}.{} WHERE {} = ? AND {} = ?",

				                    get_auth_ks_name(_qp),

				                    PERMISSIONS_CF,

				                    ROLE_NAME,

									
										15

auth/maintenance_socket_role_manager.cc
									
												View File
												
				@@ -11,6 +11,7 @@

				#include <seastar/core/future.hh>

				#include <stdexcept>

				#include <string_view>

				#include "cql3/description.hh"

				#include "utils/class_registrator.hh"

				namespace auth {

				@@ -43,10 +44,14 @@ future<> maintenance_socket_role_manager::stop() {

				    return make_ready_future<>();

				}

				future<> maintenance_socket_role_manager::ensure_superuser_is_created() {

				    return make_ready_future<>();

				}

				template<typename T = void>

				future<T> operation_not_supported_exception(std::string_view operation) {

				    return make_exception_future<T>(

				        std::runtime_error(format("role manager: {} operation not supported through maintenance socket", operation)));

				        std::runtime_error(fmt::format("role manager: {} operation not supported through maintenance socket", operation)));

				}

				future<> maintenance_socket_role_manager::create(std::string_view role_name, const role_config&, ::service::group0_batch&) {

				@@ -73,6 +78,10 @@ future<role_set> maintenance_socket_role_manager::query_granted(std::string_view

				    return operation_not_supported_exception<role_set>("QUERY GRANTED");

				}

				future<role_to_directly_granted_map> maintenance_socket_role_manager::query_all_directly_granted() {

				    return operation_not_supported_exception<role_to_directly_granted_map>("QUERY ALL DIRECTLY GRANTED");

				}

				future<role_set> maintenance_socket_role_manager::query_all() {

				    return operation_not_supported_exception<role_set>("QUERY ALL");

				}

				@@ -105,4 +114,8 @@ future<> maintenance_socket_role_manager::remove_attribute(std::string_view role

				    return operation_not_supported_exception("REMOVE ATTRIBUTE");

				}

				future<std::vector<cql3::description>> maintenance_socket_role_manager::describe_role_grants() {

				    return operation_not_supported_exception<std::vector<cql3::description>>("DESCRIBE SCHEMA WITH INTERNALS");

				}

				} // namespace auth

									
										6

auth/maintenance_socket_role_manager.hh
									
												View File
												
				@@ -39,6 +39,8 @@ public:

				    virtual future<> stop() override;

				    virtual future<> ensure_superuser_is_created() override;

				    virtual future<> create(std::string_view role_name, const role_config&, ::service::group0_batch&) override;

				    virtual future<> drop(std::string_view role_name, ::service::group0_batch& mc) override;

				@@ -51,6 +53,8 @@ public:

				    virtual future<role_set> query_granted(std::string_view grantee_name, recursive_role_query) override;

				    virtual future<role_to_directly_granted_map> query_all_directly_granted() override;

				    virtual future<role_set> query_all() override;

				    virtual future<bool> exists(std::string_view role_name) override;

				@@ -66,6 +70,8 @@ public:

				    virtual future<> set_attribute(std::string_view role_name, std::string_view attribute_name, std::string_view attribute_value, ::service::group0_batch& mc) override;

				    virtual future<> remove_attribute(std::string_view role_name, std::string_view attribute_name, ::service::group0_batch& mc) override;

				    virtual future<std::vector<cql3::description>> describe_role_grants() override;

				};

				}

									
										100

auth/password_authenticator.cc
									
												View File
												
				@@ -14,16 +14,17 @@

				#include <string_view>

				#include <optional>

				#include <boost/algorithm/cxx11/all_of.hpp>

				#include <seastar/core/seastar.hh>

				#include <seastar/core/sleep.hh>

				#include <variant>

				#include "auth/authenticated_user.hh"

				#include "auth/authentication_options.hh"

				#include "auth/common.hh"

				#include "auth/passwords.hh"

				#include "auth/roles-metadata.hh"

				#include "cql3/untyped_result_set.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "service/migration_manager.hh"

				#include "utils/class_registrator.hh"

				#include "replica/database.hh"

				@@ -75,7 +76,7 @@ static bool has_salted_hash(const cql3::untyped_result_set_row& row) {

				}

				sstring password_authenticator::update_row_query() const {

				    return format("UPDATE {}.{} SET {} = ? WHERE {} = ?",

				    return seastar::format("UPDATE {}.{} SET {} = ? WHERE {} = ?",

				            get_auth_ks_name(_qp),

				            meta::roles_table::name,

				            SALTED_HASH,

				@@ -90,7 +91,7 @@ bool password_authenticator::legacy_metadata_exists() const {

				future<> password_authenticator::migrate_legacy_metadata() const {

				    plogger.info("Starting migration of legacy authentication metadata.");

				    static const sstring query = format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, legacy_table_name);

				    static const sstring query = seastar::format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, legacy_table_name);

				    return _qp.execute_internal(

				            query,

				@@ -136,7 +137,7 @@ future<> password_authenticator::create_default_if_missing() {

				        plogger.info("Created default superuser authentication record.");

				    } else {

				        co_await announce_mutations(_qp, _group0_client, query,

				            {salted_pwd, _superuser}, &_as, ::service::raft_timeout{});

				            {salted_pwd, _superuser}, _as, ::service::raft_timeout{});

				        plogger.info("Created default superuser authentication record.");

				    }

				}

				@@ -199,7 +200,7 @@ bool password_authenticator::require_authentication() const {

				}

				authentication_option_set password_authenticator::supported_options() const {

				    return authentication_option_set{authentication_option::password};

				    return authentication_option_set{authentication_option::password, authentication_option::hashed_password};

				}

				authentication_option_set password_authenticator::alterable_options() const {

				@@ -218,28 +219,8 @@ future<authenticated_user> password_authenticator::authenticate(

				    const sstring username = credentials.at(USERNAME_KEY);

				    const sstring password = credentials.at(PASSWORD_KEY);

				    // Here was a thread local, explicit cache of prepared statement. In normal execution this is

				    // fine, but since we in testing set up and tear down system over and over, we'd start using

				    // obsolete prepared statements pretty quickly.

				    // Rely on query processing caching statements instead, and lets assume

				    // that a map lookup string->statement is not gonna kill us much.

				    const sstring query = format("SELECT {} FROM {}.{} WHERE {} = ?",

				                SALTED_HASH,

				                get_auth_ks_name(_qp),

				                meta::roles_table::name,

				                meta::roles_table::role_col_name);

				    try {

				        const auto res = co_await _qp.execute_internal(

				                query,

				                consistency_for_user(username),

				                internal_distributed_query_state(),

				                {username},

				                cql3::query_processor::cache_internal::yes);

				        auto salted_hash = std::optional<sstring>();

				        if (!res->empty()) {

				            salted_hash = res->one().get_opt<sstring>(SALTED_HASH);

				        }

				        const std::optional<sstring> salted_hash = co_await get_password_hash(username);

				        if (!salted_hash || !passwords::check(password, *salted_hash)) {

				            throw exceptions::authentication_exception("Username and/or password are incorrect");

				        }

				@@ -258,29 +239,46 @@ future<authenticated_user> password_authenticator::authenticate(

				}

				future<> password_authenticator::create(std::string_view role_name, const authentication_options& options, ::service::group0_batch& mc) {

				    if (!options.password) {

				    // When creating a role with the usual `CREATE ROLE` statement, turns the underlying `PASSWORD`

				    // into the corresponding hash.

				    // When creating a role with `CREATE ROLE WITH HASHED PASSWORD`, simply extracts the `HASHED PASSWORD`.

				    auto maybe_hash = options.credentials.transform([&] (const auto& creds) -> sstring {

				        return std::visit(make_visitor(

				                [&] (const password_option& opt) {

				                    return passwords::hash(opt.password, rng_for_salt);

				                },

				                [] (const hashed_password_option& opt) {

				                    return opt.hashed_password;

				                }

				        ), creds);

				    });

				    // Neither `PASSWORD`, nor `HASHED PASSWORD` has been specified.

				    if (!maybe_hash) {

				        co_return;

				    }

				    const auto query = update_row_query();

				    if (legacy_mode(_qp)) {

				        co_await _qp.execute_internal(

				                query,

				                consistency_for_user(role_name),

				                internal_distributed_query_state(),

				                {passwords::hash(*options.password, rng_for_salt), sstring(role_name)},

				                {std::move(*maybe_hash), sstring(role_name)},

				                cql3::query_processor::cache_internal::no).discard_result();

				    } else {

				        co_await collect_mutations(_qp, mc, query,

				                {passwords::hash(*options.password, rng_for_salt), sstring(role_name)});

				        co_await collect_mutations(_qp, mc, query, {std::move(*maybe_hash), sstring(role_name)});

				    }

				}

				future<> password_authenticator::alter(std::string_view role_name, const authentication_options& options, ::service::group0_batch& mc) {

				    if (!options.password) {

				    if (!options.credentials) {

				        co_return;

				    }

				    const sstring query = format("UPDATE {}.{} SET {} = ? WHERE {} = ?",

				    const auto password = std::get<password_option>(*options.credentials).password;

				    const sstring query = seastar::format("UPDATE {}.{} SET {} = ? WHERE {} = ?",

				            get_auth_ks_name(_qp),

				            meta::roles_table::name,

				            SALTED_HASH,

				@@ -290,16 +288,16 @@ future<> password_authenticator::alter(std::string_view role_name, const authent

				                query,

				                consistency_for_user(role_name),

				                internal_distributed_query_state(),

				                {passwords::hash(*options.password, rng_for_salt), sstring(role_name)},

				                {passwords::hash(password, rng_for_salt), sstring(role_name)},

				                cql3::query_processor::cache_internal::no).discard_result();

				    } else {

				        co_await collect_mutations(_qp, mc, query,

				                {passwords::hash(*options.password, rng_for_salt), sstring(role_name)});

				                {passwords::hash(password, rng_for_salt), sstring(role_name)});

				    }

				}

				future<> password_authenticator::drop(std::string_view name, ::service::group0_batch& mc) {

				    const sstring query = format("DELETE {} FROM {}.{} WHERE {} = ?",

				    const sstring query = seastar::format("DELETE {} FROM {}.{} WHERE {} = ?",

				            SALTED_HASH,

				            get_auth_ks_name(_qp),

				            meta::roles_table::name,

				@@ -319,6 +317,36 @@ future<custom_options> password_authenticator::query_custom_options(std::string_

				    return make_ready_future<custom_options>();

				}

				bool password_authenticator::uses_password_hashes() const {

				    return true;

				}

				future<std::optional<sstring>> password_authenticator::get_password_hash(std::string_view role_name) const {

				    // Here was a thread local, explicit cache of prepared statement. In normal execution this is

				    // fine, but since we in testing set up and tear down system over and over, we'd start using

				    // obsolete prepared statements pretty quickly.

				    // Rely on query processing caching statements instead, and lets assume

				    // that a map lookup string->statement is not gonna kill us much.

				    const sstring query = seastar::format("SELECT {} FROM {}.{} WHERE {} = ?",

				                SALTED_HASH,

				                get_auth_ks_name(_qp),

				                meta::roles_table::name,

				                meta::roles_table::role_col_name);

				    const auto res = co_await _qp.execute_internal(

				            query,

				            consistency_for_user(role_name),

				            internal_distributed_query_state(),

				            {role_name},

				            cql3::query_processor::cache_internal::yes);

				    if (res->empty()) {

				        co_return std::nullopt;

				    }

				    co_return res->one().get_opt<sstring>(SALTED_HASH);

				}

				const resource_set& password_authenticator::protected_resources() const {

				    static const resource_set resources({make_data_resource(meta::legacy::AUTH_KS, meta::roles_table::name)});

				    return resources;

									
										4

auth/password_authenticator.hh
									
												View File
												
				@@ -72,6 +72,10 @@ public:

				    virtual future<custom_options> query_custom_options(std::string_view role_name) const override;

				    virtual bool uses_password_hashes() const override;

				    virtual future<std::optional<sstring>> get_password_hash(std::string_view role_name) const override;

				    virtual const resource_set& protected_resources() const override;

				    virtual ::shared_ptr<sasl_challenge> new_sasl_challenge() const override;

									
										2

auth/permissions_cache.hh
									
												View File
												
				@@ -17,7 +17,7 @@

				#include "auth/permission.hh"

				#include "auth/resource.hh"

				#include "auth/role_or_anonymous.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "utils/hash.hh"

				#include "utils/loading_cache.hh"

									
										6

auth/resource.cc
									
												View File
												
				@@ -23,7 +23,7 @@

				#include "cql3/functions/user_function.hh"

				#include "cql3/util.hh"

				#include "db/marshal/type_parser.hh"

				#include "log.hh"

				#include "utils/log.hh"

				namespace auth {

				@@ -193,7 +193,7 @@ service_level_resource_view::service_level_resource_view(const resource &r) {

				}

				sstring encode_signature(std::string_view name, std::vector<data_type> args) {

				    return format("{}[{}]", name,

				    return seastar::format("{}[{}]", name,

				            fmt::join(args | boost::adaptors::transformed([] (const data_type t) {

				                return t->name();

				            }), "^"));

				@@ -222,7 +222,7 @@ std::pair<sstring, std::vector<data_type>> decode_signature(std::string_view enc

				// to the short form (int)

				static sstring decoded_signature_string(std::string_view encoded_signature) {

				    auto [function_name, arg_types] = decode_signature(encoded_signature);

				    return format("{}({})", cql3::util::maybe_quote(sstring(function_name)),

				    return seastar::format("{}({})", cql3::util::maybe_quote(sstring(function_name)),

				            boost::algorithm::join(arg_types | boost::adaptors::transformed([] (data_type t) {

				                return t->cql3_type_name();

				            }), ", "));

									
										5

auth/resource.hh
									
												View File
												
				@@ -18,7 +18,6 @@

				#include <unordered_set>

				#include <fmt/core.h>

				#include <seastar/core/print.hh>

				#include <seastar/core/sstring.hh>

				#include "auth/permission.hh"

				@@ -33,7 +32,7 @@ namespace auth {

				class invalid_resource_name : public std::invalid_argument {

				public:

				    explicit invalid_resource_name(std::string_view name)

				            : std::invalid_argument(format("The resource name '{}' is invalid.", name)) {

				            : std::invalid_argument(fmt::format("The resource name '{}' is invalid.", name)) {

				    }

				};

				@@ -149,7 +148,7 @@ class resource_kind_mismatch : public std::invalid_argument {

				public:

				    explicit resource_kind_mismatch(resource_kind expected, resource_kind actual)

				        : std::invalid_argument(

				            format("This resource has kind '{}', but was expected to have kind '{}'.", actual, expected)) {

				            fmt::format("This resource has kind '{}', but was expected to have kind '{}'.", actual, expected)) {

				    }

				};

									
										36

auth/role_manager.hh
									
												View File
												
				@@ -18,6 +18,7 @@

				#include <seastar/core/sstring.hh>

				#include "auth/resource.hh"

				#include "cql3/description.hh"

				#include "seastarx.hh"

				#include "exceptions/exceptions.hh"

				#include "service/raft/raft_group0_client.hh"

				@@ -48,14 +49,14 @@ public:

				class role_already_exists : public roles_argument_exception {

				public:

				    explicit role_already_exists(std::string_view role_name)

				            : roles_argument_exception(format("Role {} already exists.", role_name)) {

				            : roles_argument_exception(seastar::format("Role {} already exists.", role_name)) {

				    }

				};

				class nonexistant_role : public roles_argument_exception {

				public:

				    explicit nonexistant_role(std::string_view role_name)

				            : roles_argument_exception(format("Role {} doesn't exist.", role_name)) {

				            : roles_argument_exception(seastar::format("Role {} doesn't exist.", role_name)) {

				    }

				};

				@@ -63,7 +64,7 @@ class role_already_included : public roles_argument_exception {

				public:

				    role_already_included(std::string_view grantee_name, std::string_view role_name)

				            : roles_argument_exception(

				                      format("{} already includes role {}.", grantee_name, role_name)) {

				                      seastar::format("{} already includes role {}.", grantee_name, role_name)) {

				    }

				};

				@@ -71,11 +72,12 @@ class revoke_ungranted_role : public roles_argument_exception {

				public:

				    revoke_ungranted_role(std::string_view revokee_name, std::string_view role_name)

				            : roles_argument_exception(

				                      format("{} was not granted role {}, so it cannot be revoked.", revokee_name, role_name)) {

				                      seastar::format("{} was not granted role {}, so it cannot be revoked.", revokee_name, role_name)) {

				    }

				};

				using role_set = std::unordered_set<sstring>;

				using role_to_directly_granted_map = std::multimap<sstring, sstring>;

				enum class recursive_role_query { yes, no };

				@@ -105,6 +107,13 @@ public:

				    virtual future<> stop() = 0;

				    ///

				    /// Ensure that superuser role exists.

				    ///

				    /// \returns a future once it is ensured that the superuser role exists.

				    ///

				    virtual future<> ensure_superuser_is_created() = 0;

				    ///

				    /// \returns an exceptional future with \ref role_already_exists for a role that has previously been created.

				    ///

				@@ -144,6 +153,22 @@ public:

				    ///

				    virtual future<role_set> query_granted(std::string_view grantee, recursive_role_query) = 0;

				    /// \returns map of directly granted roles for all roles

				    ///

				    /// Example:

				    /// GRANT role2 TO role1

				    /// GRANT role3 TO role1

				    /// GRANT role3 TO role2

				    ///

				    /// Will return map:

				    /// {

				    ///   (role1, role2),

				    ///   (role1, role3),

				    ///   (role2, role3)

				    /// }

				    ///  

				    virtual future<role_to_directly_granted_map> query_all_directly_granted() = 0;

				    virtual future<role_set> query_all() = 0;

				    virtual future<bool> exists(std::string_view role_name) = 0;

				@@ -178,5 +203,8 @@ public:

				    /// \note: This is a no-op if the role does not have the named attribute set.

				    ///

				    virtual future<> remove_attribute(std::string_view role_name, std::string_view attribute_name, ::service::group0_batch& mc) = 0;

				    /// Produces descriptions that can be used to restore the role grants.

				    virtual future<std::vector<cql3::description>> describe_role_grants() = 0;

				};

				}

									
										7

auth/roles-metadata.cc
									
												View File
												
				@@ -8,7 +8,6 @@

				#include "auth/roles-metadata.hh"

				#include <boost/algorithm/cxx11/any_of.hpp>

				#include <seastar/core/print.hh>

				#include <seastar/core/shared_ptr.hh>

				#include <seastar/core/sstring.hh>

				@@ -47,7 +46,7 @@ future<bool> default_role_row_satisfies(

				        cql3::query_processor& qp,

				        std::function<bool(const cql3::untyped_result_set_row&)> p,

				        std::optional<std::string> rolename) {

				    const sstring query = format("SELECT * FROM {}.{} WHERE {} = ?",

				    const sstring query = seastar::format("SELECT * FROM {}.{} WHERE {} = ?",

				            get_auth_ks_name(qp),

				            meta::roles_table::name,

				            meta::roles_table::role_col_name);

				@@ -69,7 +68,7 @@ future<bool> any_nondefault_role_row_satisfies(

				        cql3::query_processor& qp,

				        std::function<bool(const cql3::untyped_result_set_row&)> p,

				        std::optional<std::string> rolename) {

				    const sstring query = format("SELECT * FROM {}.{}", get_auth_ks_name(qp), meta::roles_table::name);

				    const sstring query = seastar::format("SELECT * FROM {}.{}", get_auth_ks_name(qp), meta::roles_table::name);

				    auto results = co_await qp.execute_internal(query, db::consistency_level::QUORUM

				        , internal_distributed_query_state(), cql3::query_processor::cache_internal::no

				@@ -79,7 +78,7 @@ future<bool> any_nondefault_role_row_satisfies(

				    }

				    static const sstring col_name = sstring(meta::roles_table::role_col_name);

				    co_return boost::algorithm::any_of(*results, [&](const cql3::untyped_result_set_row& row) {

				    co_return std::ranges::any_of(*results, [&](const cql3::untyped_result_set_row& row) {

				        auto superuser = rolename ? std::string_view(*rolename) : meta::DEFAULT_SUPERUSER_NAME;

				        const bool is_nondefault = row.get_as<sstring>(col_name) != superuser;

				        return is_nondefault && p(row);

									
										2

auth/sasl_challenge.hh
									
												View File
												
				@@ -18,7 +18,7 @@

				#include <seastar/core/sstring.hh>

				#include "auth/authenticated_user.hh"

				#include "bytes.hh"

				#include "bytes_fwd.hh"

				#include "seastarx.hh"

				namespace auth {

									
										246

auth/service.cc
									
												View File
												
				@@ -8,6 +8,8 @@

				#include <exception>

				#include <seastar/core/coroutine.hh>

				#include "auth/authentication_options.hh"

				#include "auth/authorizer.hh"

				#include "auth/resource.hh"

				#include "auth/service.hh"

				@@ -25,17 +27,21 @@

				#include "auth/role_or_anonymous.hh"

				#include "cql3/functions/functions.hh"

				#include "cql3/query_processor.hh"

				#include "cql3/description.hh"

				#include "cql3/untyped_result_set.hh"

				#include "cql3/util.hh"

				#include "db/config.hh"

				#include "db/consistency_level_type.hh"

				#include "db/functions/function_name.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "schema/schema_fwd.hh"

				#include <seastar/core/future.hh>

				#include <seastar/coroutine/parallel_for_each.hh>

				#include <variant>

				#include "service/migration_manager.hh"

				#include "service/raft/raft_group0_client.hh"

				#include "timestamp.hh"

				#include "utils/assert.hh"

				#include "utils/class_registrator.hh"

				#include "locator/abstract_replication_strategy.hh"

				#include "data_dictionary/keyspace_metadata.hh"

				@@ -77,7 +83,7 @@ private:

				    void on_update_function(const sstring& ks_name, const sstring& function_name) override {}

				    void on_update_aggregate(const sstring& ks_name, const sstring& aggregate_name) override {}

				    void on_update_view(const sstring& ks_name, const sstring& view_name, bool columns_changed) override {}

				    void on_update_tablet_metadata() override {}

				    void on_update_tablet_metadata(const locator::tablet_metadata_change_hint&) override {}

				    void on_drop_keyspace(const sstring& ks_name) override {

				        if (!legacy_mode(_qp)) {

				@@ -194,7 +200,7 @@ service::service(

				}

				future<> service::create_legacy_keyspace_if_missing(::service::migration_manager& mm) const {

				    assert(this_shard_id() == 0); // once_among_shards makes sure a function is executed on shard 0 only

				    SCYLLA_ASSERT(this_shard_id() == 0); // once_among_shards makes sure a function is executed on shard 0 only

				    auto db = _qp.db();

				    while (!db.has_keyspace(meta::legacy::AUTH_KS)) {

				@@ -212,7 +218,7 @@ future<> service::create_legacy_keyspace_if_missing(::service::migration_manager

				            try {

				                co_return co_await mm.announce(::service::prepare_new_keyspace_announcement(db.real_database(), ksm, ts),

				                        std::move(group0_guard), format("auth_service: create {} keyspace", meta::legacy::AUTH_KS));

				                        std::move(group0_guard), seastar::format("auth_service: create {} keyspace", meta::legacy::AUTH_KS));

				            } catch (::service::group0_concurrent_modification&) {

				                log.info("Concurrent operation is detected while creating {} keyspace, retrying.", meta::legacy::AUTH_KS);

				            }

				@@ -256,6 +262,10 @@ future<> service::stop() {

				    });

				}

				future<> service::ensure_superuser_is_created() {

				    return _role_manager->ensure_superuser_is_created();

				}

				void service::update_cache_config() {

				    auto db = _qp.db();

				@@ -321,8 +331,11 @@ static void validate_authentication_options_are_supported(

				        }

				    };

				    if (options.password) {

				        check(authentication_option::password);

				    if (options.credentials) {

				        std::visit(make_visitor(

				            [&] (const password_option&) { check(authentication_option::password); },

				            [&] (const hashed_password_option&) { check(authentication_option::hashed_password); }

				        ), *options.credentials);

				    }

				    if (options.options) {

				@@ -417,6 +430,223 @@ future<bool> service::exists(const resource& r) const {

				    return make_ready_future<bool>(false);

				}

				future<std::vector<cql3::description>> service::describe_roles(bool with_hashed_passwords) {

				    std::vector<cql3::description> result{};

				    const auto roles = co_await _role_manager->query_all();

				    result.reserve(roles.size());

				    const bool authenticator_uses_password_hashes = _authenticator->uses_password_hashes();

				    auto produce_create_statement = [with_hashed_passwords] (const sstring& formatted_role_name,

				            const std::optional<sstring>& maybe_hashed_password, bool can_login, bool is_superuser) {

				        // Even after applying formatting to a role, `formatted_role_name` can only equal `meta::DEFAULT_SUPER_NAME`

				        // if the original identifier was equal to it.

				        const sstring role_part = formatted_role_name == meta::DEFAULT_SUPERUSER_NAME

				                ? seastar::format("IF NOT EXISTS {}", formatted_role_name)

				                : formatted_role_name;

				        const sstring with_hashed_password_part = with_hashed_passwords && maybe_hashed_password

				                // `K_PASSWORD` in Scylla's CQL grammar requires that passwords be quoted

				                // with single quotation marks.

				                ? seastar::format("WITH HASHED PASSWORD = {} AND", cql3::util::single_quote(*maybe_hashed_password))

				                : "WITH";

				        return seastar::format("CREATE ROLE {} {} LOGIN = {} AND SUPERUSER = {};",

				                role_part, with_hashed_password_part, can_login, is_superuser);

				    };

				    for (const auto& role : roles) {

				        const sstring formatted_role_name = cql3::util::maybe_quote(role);

				        std::optional<sstring> maybe_hashed_password;

				        if (authenticator_uses_password_hashes) {

				            maybe_hashed_password = co_await _authenticator->get_password_hash(role);

				        }

				        const bool can_login = co_await _role_manager->can_login(role);

				        const bool is_superuser = co_await _role_manager->is_superuser(role);

				        result.push_back(cql3::description {

				            // Roles do not belong to any keyspace.

				            .keyspace = std::nullopt,

				            .type = "role",

				            .name = role,

				            .create_statement = produce_create_statement(formatted_role_name, maybe_hashed_password, can_login, is_superuser)

				        });

				    }

				    std::ranges::sort(result, std::less<>{}, std::mem_fn(&cql3::description::name));

				    co_return result;

				}

				// The function doesn't assume anything about `role`.

				static sstring describe_data_resource(const permission& perm, const resource& r, std::string_view role) {

				    const auto permission = permissions::to_string(perm);

				    const auto formatted_role = cql3::util::maybe_quote(role);

				    const auto view = data_resource_view(r);

				    const auto maybe_ks = view.keyspace();

				    const auto maybe_cf = view.table();

				    // The documentation says:

				    //

				    //     Both keyspace and table names consist of only alphanumeric characters, cannot be empty,

				    //     and are limited in size to 48 characters (that limit exists mostly to avoid filenames,

				    //     which may include the keyspace and table name, to go over the limits of certain file systems).

				    //     By default, keyspace and table names are case insensitive (myTable is equivalent to mytable),

				    //     but case sensitivity can be forced by using double-quotes ("myTable" is different from mytable).

				    //

				    // That's why we wrap identifiers with quotation marks below.

				    if (!maybe_ks) {

				        return seastar::format("GRANT {} ON ALL KEYSPACES TO {};", permission, formatted_role);

				    }

				    const auto ks = cql3::util::maybe_quote(*maybe_ks);

				    if (!maybe_cf) {

				        return seastar::format("GRANT {} ON KEYSPACE {} TO {};", permission, ks, formatted_role);

				    }

				    const auto cf = cql3::util::maybe_quote(*maybe_cf);

				    return seastar::format("GRANT {} ON {}.{} TO {};", permission, ks, cf, formatted_role);

				}

				// The function doesn't assume anything about `role`.

				static sstring describe_role_resource(const permission& perm, const resource& r, std::string_view role) {

				    const auto permission = permissions::to_string(perm);

				    const auto formatted_role = cql3::util::maybe_quote(role);

				    const auto view = role_resource_view(r);

				    const auto maybe_target_role = view.role();

				    if (!maybe_target_role) {

				        return seastar::format("GRANT {} ON ALL ROLES TO {};", permission, formatted_role);

				    }

				    return seastar::format("GRANT {} ON ROLE {} TO {};", permission, cql3::util::maybe_quote(*maybe_target_role), formatted_role);

				}

				// The function doesn't assume anything about `role`.

				static sstring describe_udf_resource(const permission& perm, const resource& r, std::string_view role) {

				    const auto permission = permissions::to_string(perm);

				    const auto formatted_role = cql3::util::maybe_quote(role);

				    const auto view = functions_resource_view(r);

				    const auto maybe_ks = view.keyspace();

				    const auto maybe_fun_sig = view.function_signature();

				    const auto maybe_fun_name = view.function_name();

				    const auto maybe_fun_args = view.function_args();

				    // The documentation says:

				    //

				    //     Both keyspace and table names consist of only alphanumeric characters, cannot be empty,

				    //     and are limited in size to 48 characters (that limit exists mostly to avoid filenames,

				    //     which may include the keyspace and table name, to go over the limits of certain file systems).

				    //     By default, keyspace and table names are case insensitive (myTable is equivalent to mytable),

				    //     but case sensitivity can be forced by using double-quotes ("myTable" is different from mytable).

				    //

				    // That's why we wrap identifiers with quotation marks below.

				    if (!maybe_ks) {

				        return seastar::format("GRANT {} ON ALL FUNCTIONS TO {};", permission, formatted_role);

				    }

				    const auto ks = cql3::util::maybe_quote(*maybe_ks);

				    if (!maybe_fun_sig && !maybe_fun_name) {

				        return seastar::format("GRANT {} ON ALL FUNCTIONS IN KEYSPACE {} TO {};", permission, ks, formatted_role);

				    }

				    if (maybe_fun_name) {

				        SCYLLA_ASSERT(maybe_fun_args);

				        const auto fun_name = cql3::util::maybe_quote(*maybe_fun_name);

				        const auto fun_args_range = *maybe_fun_args | std::views::transform([] (const auto& fun_arg) {

				            return cql3::util::maybe_quote(fun_arg);

				        });

				        return seastar::format("GRANT {} ON FUNCTION {}.{}({}) TO {};",

				                permission, ks, fun_name, fmt::join(fun_args_range, ", "), formatted_role);

				    }

				    SCYLLA_ASSERT(maybe_fun_sig);

				    auto [fun_name, fun_args] = decode_signature(*maybe_fun_sig);

				    fun_name = cql3::util::maybe_quote(fun_name);

				    // We don't call `cql3::util::maybe_quote` later because `cql3_type_name_without_frozen` already guarantees

				    // that the type will be wrapped within double quotation marks if it's necessary.

				    auto parsed_fun_args = fun_args | std::views::transform([] (const data_type& dt) {

				        return dt->without_reversed().cql3_type_name_without_frozen();

				    });

				    return seastar::format("GRANT {} ON FUNCTION {}.{}({}) TO {};",

				            permission, ks, fun_name, fmt::join(parsed_fun_args, ", "), formatted_role);

				}

				// The function doesn't assume anything about `role`.

				static sstring describe_resource_kind(const permission& perm, const resource& r, std::string_view role) {

				    switch (r.kind()) {

				        case resource_kind::data:

				            return describe_data_resource(perm, r, role);

				        case resource_kind::role:

				            return describe_role_resource(perm, r, role);

				        case resource_kind::service_level:

				            on_internal_error(log, "Granting permissions for service levels is not supported");

				        case resource_kind::functions:

				            return describe_udf_resource(perm, r, role);

				    }

				}

				future<std::vector<cql3::description>> service::describe_permissions() const {

				    std::vector<cql3::description> result{};

				    const auto permission_list = co_await std::invoke([&] -> future<std::vector<permission_details>> {

				        try {

				            co_return co_await _authorizer->list_all();

				        } catch (const unsupported_authorization_operation&) {

				            // If Scylla uses AllowAllAuthorizer, permissions do not exist and the corresponding authorizer

				            // will throw an exception when trying to access them.

				            co_return std::vector<permission_details>{};

				        }

				    });

				    for (const auto& permissions : permission_list) {

				        for (const auto& permission : permissions.permissions) {

				            result.push_back(cql3::description {

				                // Permission grants do not belong to any keyspace.

				                .keyspace = std::nullopt,

				                .type = "grant_permission",

				                .name = permissions.role_name,

				                .create_statement = describe_resource_kind(permission, permissions.resource, permissions.role_name)

				            });

				        }

				        co_await coroutine::maybe_yield();

				    }

				    std::ranges::sort(result, std::less<>{}, [] (const cql3::description& desc) noexcept {

				        return std::make_tuple(std::ref(desc.name), std::ref(*desc.create_statement));

				    });

				    co_return result;

				}

				future<std::vector<cql3::description>> service::describe_auth(bool with_hashed_passwords) {

				    auto role_descs = co_await describe_roles(with_hashed_passwords);

				    auto role_grant_descs = co_await _role_manager->describe_role_grants();

				    auto permission_descs = co_await describe_permissions();

				    auto join_vectors = [] (std::vector<cql3::description>& v1, std::vector<cql3::description>&& v2) {

				        v1.insert(v1.end(), std::make_move_iterator(v2.begin()), std::make_move_iterator(v2.end()));

				    };

				    join_vectors(role_descs, std::move(role_grant_descs));

				    join_vectors(role_descs, std::move(permission_descs));

				    co_return role_descs;

				}

				//

				// Free functions.

				//

				@@ -632,7 +862,7 @@ future<> migrate_to_auth_v2(db::system_keyspace& sys_ks, ::service::raft_group0_

				            ::service::query_state qs(cs, empty_service_permit());

				            auto rows = co_await qp.execute_internal(

				                    format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, cf_name),

				                    seastar::format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, cf_name),

				                    db::consistency_level::ALL,

				                    qs,

				                    {},

				@@ -681,7 +911,7 @@ future<> migrate_to_auth_v2(db::system_keyspace& sys_ks, ::service::raft_group0_

				    co_await announce_mutations_with_batching(g0,

				            start_operation_func,

				            std::move(gen),

				            &as,

				            as,

				            std::nullopt);

				}

									
										12

auth/service.hh
									
												View File
												
				@@ -23,6 +23,7 @@

				#include "auth/permissions_cache.hh"

				#include "auth/role_manager.hh"

				#include "auth/common.hh"

				#include "cql3/description.hh"

				#include "seastarx.hh"

				#include "service/raft/raft_group0_client.hh"

				#include "utils/observable.hh"

				@@ -131,6 +132,8 @@ public:

				    future<> stop();

				    future<> ensure_superuser_is_created();

				    void update_cache_config();

				    void reset_authorization_cache();

				@@ -174,6 +177,12 @@ public:

				    future<bool> exists(const resource&) const;

				    ///

				    /// Produces descriptions that can be used to restore the state of auth. That encompasses

				    /// roles, role grants, and permission grants.

				    ///

				    future<std::vector<cql3::description>> describe_auth(bool with_hashed_passwords);

				    authenticator& underlying_authenticator() const {

				        return *_authenticator;

				    }

				@@ -197,6 +206,9 @@ public:

				private:

				    future<> create_legacy_keyspace_if_missing(::service::migration_manager& mm) const;

				    future<bool> has_superuser(std::string_view role_name, const role_set& roles) const;

				    future<std::vector<cql3::description>> describe_roles(bool with_hashed_passwords);

				    future<std::vector<cql3::description>> describe_permissions() const;

				};

				future<bool> has_superuser(const service&, const authenticated_user&);

									
										151

auth/standard_role_manager.cc
									
												View File
												
				@@ -21,12 +21,17 @@

				#include <seastar/core/thread.hh>

				#include "auth/common.hh"

				#include "auth/role_manager.hh"

				#include "auth/roles-metadata.hh"

				#include "cql3/query_processor.hh"

				#include "cql3/description.hh"

				#include "cql3/untyped_result_set.hh"

				#include "cql3/util.hh"

				#include "db/consistency_level_type.hh"

				#include "exceptions/exceptions.hh"

				#include "log.hh"

				#include "utils/log.hh"

				#include "seastar/core/loop.hh"

				#include "seastar/coroutine/maybe_yield.hh"

				#include "service/raft/raft_group0_client.hh"

				#include "utils/class_registrator.hh"

				#include "service/migration_manager.hh"

				@@ -46,7 +51,7 @@ namespace role_attributes_table {

				constexpr std::string_view name{"role_attributes", 15};

				static std::string_view creation_query() noexcept {

				    static const sstring instance = format(

				    static const sstring instance = seastar::format(

				            "CREATE TABLE {}.{} ("

				            "  role text,"

				            "  name text,"

				@@ -86,7 +91,7 @@ static db::consistency_level consistency_for_role(std::string_view role_name) no

				}

				static future<std::optional<record>> find_record(cql3::query_processor& qp, std::string_view role_name) {

				    const sstring query = format("SELECT * FROM {}.{} WHERE {} = ?",

				    const sstring query = seastar::format("SELECT * FROM {}.{} WHERE {} = ?",

				            get_auth_ks_name(qp),

				            meta::roles_table::name,

				            meta::roles_table::role_col_name);

				@@ -180,7 +185,7 @@ future<> standard_role_manager::create_default_role_if_missing() {

				        if (exists) {

				            co_return;

				        }

				        const sstring query = format("INSERT INTO {}.{} ({}, is_superuser, can_login) VALUES (?, true, true)",

				        const sstring query = seastar::format("INSERT INTO {}.{} ({}, is_superuser, can_login) VALUES (?, true, true)",

				                get_auth_ks_name(_qp),

				                meta::roles_table::name,

				                meta::roles_table::role_col_name);

				@@ -192,7 +197,7 @@ future<> standard_role_manager::create_default_role_if_missing() {

				                    {_superuser},

				                    cql3::query_processor::cache_internal::no).discard_result();

				        } else {

				            co_await announce_mutations(_qp, _group0_client, query, {_superuser}, &_as, ::service::raft_timeout{});

				            co_await announce_mutations(_qp, _group0_client, query, {_superuser}, _as, ::service::raft_timeout{});

				        }

				        log.info("Created default superuser role '{}'.", _superuser);

				    } catch(const exceptions::unavailable_exception& e) {

				@@ -209,7 +214,7 @@ bool standard_role_manager::legacy_metadata_exists() {

				future<> standard_role_manager::migrate_legacy_metadata() {

				    log.info("Starting migration of legacy user metadata.");

				    static const sstring query = format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, legacy_table_name);

				    static const sstring query = seastar::format("SELECT * FROM {}.{}", meta::legacy::AUTH_KS, legacy_table_name);

				    return _qp.execute_internal(

				            query,

				@@ -238,35 +243,39 @@ future<> standard_role_manager::migrate_legacy_metadata() {

				}

				future<> standard_role_manager::start() {

				    return once_among_shards([this] {

				        return futurize_invoke([this] () {

				            if (legacy_mode(_qp)) {

				                return create_legacy_metadata_tables_if_missing();

				            }

				            return make_ready_future<>();

				        }).then([this] {

				            _stopped = auth::do_after_system_ready(_as, [this] {

				                return seastar::async([this] {

				                    if (legacy_mode(_qp)) {

				                        _migration_manager.wait_for_schema_agreement(_qp.db().real_database(), db::timeout_clock::time_point::max(), &_as).get();

				    return once_among_shards([this] () -> future<> {

				        if (legacy_mode(_qp)) {

				            co_await create_legacy_metadata_tables_if_missing();

				        }

				                        if (any_nondefault_role_row_satisfies(_qp, &has_can_login).get()) {

				                            if (legacy_metadata_exists()) {

				                                log.warn("Ignoring legacy user metadata since nondefault roles already exist.");

				                            }

				        auto handler = [this] () -> future<> {

				            const bool legacy = legacy_mode(_qp);

				            if (legacy) {

				                if (!_superuser_created_promise.available()) {

				                    _superuser_created_promise.set_value();

				                }

				                co_await _migration_manager.wait_for_schema_agreement(_qp.db().real_database(), db::timeout_clock::time_point::max(), &_as);

				                            return;

				                        }

				                        if (legacy_metadata_exists()) {

				                            migrate_legacy_metadata().get();

				                            return;

				                        }

				                if (co_await any_nondefault_role_row_satisfies(_qp, &has_can_login)) {

				                    if (legacy_metadata_exists()) {

				                        log.warn("Ignoring legacy user metadata since nondefault roles already exist.");

				                    }

				                    create_default_role_if_missing().get();

				                });

				            });

				        });

				                    co_return;

				                }

				                if (legacy_metadata_exists()) {

				                    co_await migrate_legacy_metadata();

				                    co_return;

				                }

				            }

				            co_await create_default_role_if_missing();

				            if (!legacy) {

				                _superuser_created_promise.set_value();

				            }

				        };

				        _stopped = auth::do_after_system_ready(_as, handler);

				        co_return;

				    });

				}

				@@ -275,8 +284,13 @@ future<> standard_role_manager::stop() {

				    return _stopped.handle_exception_type([] (const sleep_aborted&) { }).handle_exception_type([](const abort_requested_exception&) {});;

				}

				future<> standard_role_manager::ensure_superuser_is_created() {

				    SCYLLA_ASSERT(this_shard_id() == 0);

				    return _superuser_created_promise.get_shared_future();

				}

				future<> standard_role_manager::create_or_replace(std::string_view role_name, const role_config& c, ::service::group0_batch& mc) {

				    const sstring query = format("INSERT INTO {}.{} ({}, is_superuser, can_login) VALUES (?, ?, ?)",

				    const sstring query = seastar::format("INSERT INTO {}.{} ({}, is_superuser, can_login) VALUES (?, ?, ?)",

				            get_auth_ks_name(_qp),

				            meta::roles_table::name,

				            meta::roles_table::role_col_name);

				@@ -323,7 +337,7 @@ standard_role_manager::alter(std::string_view role_name, const role_config_updat

				        if (!u.is_superuser && !u.can_login) {

				            return make_ready_future<>();

				        }

				        const sstring query = format("UPDATE {}.{} SET {} WHERE {} = ?",

				        const sstring query = seastar::format("UPDATE {}.{} SET {} WHERE {} = ?",

				            get_auth_ks_name(_qp),

				            meta::roles_table::name,

				            build_column_assignments(u),

				@@ -347,7 +361,7 @@ future<> standard_role_manager::drop(std::string_view role_name, ::service::grou

				    }

				    // First, revoke this role from all roles that are members of it.

				    const auto revoke_from_members = [this, role_name, &mc] () -> future<> {

				        const sstring query = format("SELECT member FROM {}.{} WHERE role = ?",

				        const sstring query = seastar::format("SELECT member FROM {}.{} WHERE role = ?",

				                get_auth_ks_name(_qp),

				                meta::role_members_table::name);

				        const auto members = co_await _qp.execute_internal(

				@@ -379,7 +393,7 @@ future<> standard_role_manager::drop(std::string_view role_name, ::service::grou

				    };

				    // Delete all attributes for that role

				    const auto remove_attributes_of = [this, role_name, &mc] () -> future<> {

				        const sstring query = format("DELETE FROM {}.{} WHERE role = ?",

				        const sstring query = seastar::format("DELETE FROM {}.{} WHERE role = ?",

				                get_auth_ks_name(_qp),

				                meta::role_attributes_table::name);

				        if (legacy_mode(_qp)) {

				@@ -391,7 +405,7 @@ future<> standard_role_manager::drop(std::string_view role_name, ::service::grou

				    };

				    // Finally, delete the role itself.

				    const auto delete_role = [this, role_name, &mc] () -> future<> {

				        const sstring query = format("DELETE FROM {}.{} WHERE {} = ?",

				        const sstring query = seastar::format("DELETE FROM {}.{} WHERE {} = ?",

				                get_auth_ks_name(_qp),

				                meta::roles_table::name,

				                meta::roles_table::role_col_name);

				@@ -418,7 +432,7 @@ standard_role_manager::legacy_modify_membership(

				        std::string_view role_name,

				        membership_change ch) {

				    const auto modify_roles = [this, role_name, grantee_name, ch] () -> future<> {

				        const auto query = format(

				        const auto query = seastar::format(

				                "UPDATE {}.{} SET member_of = member_of {} ? WHERE {} = ?",

				                get_auth_ks_name(_qp),

				                meta::roles_table::name,

				@@ -435,7 +449,7 @@ standard_role_manager::legacy_modify_membership(

				    const auto modify_role_members = [this, role_name, grantee_name, ch] () -> future<> {

				        switch (ch) {

				            case membership_change::add: {

				                const sstring insert_query = format("INSERT INTO {}.{} (role, member) VALUES (?, ?)",

				                const sstring insert_query = seastar::format("INSERT INTO {}.{} (role, member) VALUES (?, ?)",

				                        get_auth_ks_name(_qp),

				                        meta::role_members_table::name);

				                co_return co_await _qp.execute_internal(

				@@ -447,7 +461,7 @@ standard_role_manager::legacy_modify_membership(

				            }

				            case membership_change::remove: {

				                const sstring delete_query = format("DELETE FROM {}.{} WHERE role = ? AND member = ?",

				                const sstring delete_query = seastar::format("DELETE FROM {}.{} WHERE role = ? AND member = ?",

				                        get_auth_ks_name(_qp),

				                        meta::role_members_table::name);

				                co_return co_await _qp.execute_internal(

				@@ -473,7 +487,7 @@ standard_role_manager::modify_membership(

				        co_return co_await legacy_modify_membership(grantee_name, role_name, ch);

				    }

				    const auto modify_roles = format(

				    const auto modify_roles = seastar::format(

				            "UPDATE {}.{} SET member_of = member_of {} ? WHERE {} = ?",

				            get_auth_ks_name(_qp),

				            meta::roles_table::name,

				@@ -485,12 +499,12 @@ standard_role_manager::modify_membership(

				    sstring modify_role_members;

				    switch (ch) {

				    case membership_change::add:

				        modify_role_members = format("INSERT INTO {}.{} (role, member) VALUES (?, ?)",

				        modify_role_members = seastar::format("INSERT INTO {}.{} (role, member) VALUES (?, ?)",

				                get_auth_ks_name(_qp),

				                meta::role_members_table::name);

				        break;

				    case membership_change::remove:

				        modify_role_members = format("DELETE FROM {}.{} WHERE role = ? AND member = ?",

				        modify_role_members = seastar::format("DELETE FROM {}.{} WHERE role = ? AND member = ?",

				                get_auth_ks_name(_qp),

				                meta::role_members_table::name);

				        break;

				@@ -583,8 +597,22 @@ future<role_set> standard_role_manager::query_granted(std::string_view grantee_n

				    });

				}

				future<role_to_directly_granted_map> standard_role_manager::query_all_directly_granted() {

				    const sstring query = seastar::format("SELECT * FROM {}.{}",

				            get_auth_ks_name(_qp),

				            meta::role_members_table::name);

				    role_to_directly_granted_map roles_map;

				    co_await _qp.query_internal(query, [&roles_map] (const cql3::untyped_result_set_row& row) -> future<stop_iteration> {

				        roles_map.insert({row.get_as<sstring>("member"), row.get_as<sstring>("role")});

				        co_return stop_iteration::no;

				    });

				    co_return roles_map;

				}

				future<role_set> standard_role_manager::query_all() {

				    const sstring query = format("SELECT {} FROM {}.{}",

				    const sstring query = seastar::format("SELECT {} FROM {}.{}",

				            meta::roles_table::role_col_name,

				            get_auth_ks_name(_qp),

				            meta::roles_table::name);

				@@ -628,7 +656,7 @@ future<bool> standard_role_manager::can_login(std::string_view role_name) {

				}

				future<std::optional<sstring>> standard_role_manager::get_attribute(std::string_view role_name, std::string_view attribute_name) {

				    const sstring query = format("SELECT name, value FROM {}.{} WHERE role = ? AND name = ?",

				    const sstring query = seastar::format("SELECT name, value FROM {}.{} WHERE role = ? AND name = ?",

				            get_auth_ks_name(_qp),

				            meta::role_attributes_table::name);

				    const auto result_set = co_await _qp.execute_internal(query, {sstring(role_name), sstring(attribute_name)}, cql3::query_processor::cache_internal::yes);

				@@ -659,7 +687,7 @@ future<> standard_role_manager::set_attribute(std::string_view role_name, std::s

				    if (!co_await exists(role_name)) {

				        throw auth::nonexistant_role(role_name);

				    }

				    const sstring query = format("INSERT INTO {}.{} (role, name, value)  VALUES (?, ?, ?)",

				    const sstring query = seastar::format("INSERT INTO {}.{} (role, name, value)  VALUES (?, ?, ?)",

				            get_auth_ks_name(_qp),

				            meta::role_attributes_table::name);

				    if (legacy_mode(_qp)) {

				@@ -674,7 +702,7 @@ future<> standard_role_manager::remove_attribute(std::string_view role_name, std

				    if (!co_await exists(role_name)) {

				        throw auth::nonexistant_role(role_name);

				    }

				    const sstring query = format("DELETE FROM {}.{} WHERE role = ? AND name = ?",

				    const sstring query = seastar::format("DELETE FROM {}.{} WHERE role = ? AND name = ?",

				            get_auth_ks_name(_qp),

				            meta::role_attributes_table::name);

				    if (legacy_mode(_qp)) {

				@@ -684,4 +712,33 @@ future<> standard_role_manager::remove_attribute(std::string_view role_name, std

				                {sstring(role_name), sstring(attribute_name)});

				    }

				}

				future<std::vector<cql3::description>> standard_role_manager::describe_role_grants() {

				    std::vector<cql3::description> result{};

				    const auto grants = co_await query_all_directly_granted();

				    result.reserve(grants.size());

				    for (const auto& [grantee_role, granted_role] : grants) {

				        const auto formatted_grantee = cql3::util::maybe_quote(grantee_role);

				        const auto formatted_granted = cql3::util::maybe_quote(granted_role);

				        result.push_back(cql3::description {

				            // Role grants do not belong to any keyspace.

				            .keyspace = std::nullopt,

				            .type = "grant_role",

				            .name = granted_role,

				            .create_statement = seastar::format("GRANT {} TO {};", formatted_granted, formatted_grantee)

				        });

				        co_await coroutine::maybe_yield();

				    }

				    std::ranges::sort(result, std::less<>{}, [] (const cql3::description& desc) noexcept {

				        return std::make_tuple(std::ref(desc.name), std::ref(*desc.create_statement));

				    });

				    co_return result;

				}

				} // namespace auth

									
										11

auth/standard_role_manager.hh
									
												View File
												
				@@ -15,8 +15,10 @@

				#include <seastar/core/abort_source.hh>

				#include <seastar/core/future.hh>

				#include <seastar/core/shared_future.hh>

				#include <seastar/core/sstring.hh>

				#include "cql3/description.hh"

				#include "seastarx.hh"

				#include "service/raft/raft_group0_client.hh"

				@@ -37,6 +39,7 @@ class standard_role_manager final : public role_manager {

				    future<> _stopped;

				    abort_source _as;

				    std::string _superuser;

				    shared_promise<> _superuser_created_promise;

				public:

				    standard_role_manager(cql3::query_processor&, ::service::raft_group0_client&, ::service::migration_manager&);

				@@ -49,6 +52,8 @@ public:

				    virtual future<> stop() override;

				    virtual future<> ensure_superuser_is_created() override;

				    virtual future<> create(std::string_view role_name, const role_config&, ::service::group0_batch&) override;

				    virtual future<> drop(std::string_view role_name, ::service::group0_batch& mc) override;

				@@ -61,6 +66,8 @@ public:

				    virtual future<role_set> query_granted(std::string_view grantee_name, recursive_role_query) override;

				    virtual future<role_to_directly_granted_map> query_all_directly_granted() override;

				    virtual future<role_set> query_all() override;

				    virtual future<bool> exists(std::string_view role_name) override;

				@@ -77,6 +84,8 @@ public:

				    virtual future<> remove_attribute(std::string_view role_name, std::string_view attribute_name, ::service::group0_batch& mc) override;

				    virtual future<std::vector<cql3::description>> describe_role_grants() override;

				private:

				    enum class membership_change { add, remove };

				@@ -95,4 +104,4 @@ private:

				    future<> modify_membership(std::string_view role_name, std::string_view grantee_name, membership_change, ::service::group0_batch& mc);

				};

				}

				} // namespace auth

									
										8

auth/transitional.cc
									
												View File
												
				@@ -103,6 +103,14 @@ public:

				        return _authenticator->query_custom_options(role_name);

				    }

				    virtual bool uses_password_hashes() const override {

				        return _authenticator->uses_password_hashes();

				    }

				    virtual future<std::optional<sstring>> get_password_hash(std::string_view role_name) const override {

				        return _authenticator->get_password_hash(role_name);

				    }

				    virtual const resource_set& protected_resources() const override {

				        return _authenticator->protected_resources();

				    }

2

bin/cqlsh

View File

@@ -4,5 +4,5 @@
 # SPDX-License-Identifier: AGPL-3.0-or-later
 here=$(dirname "$0")
 exec "$here/../tools/cqlsh/bin/cqlsh" "$@"
 exec "$here/../tools/cqlsh/bin/cqlsh.py" "$@"

									
										3

bytes.cc
									
												View File
												
				@@ -7,6 +7,7 @@

				 */

				#include "bytes.hh"

				#include <fmt/ostream.h>

				#include <seastar/core/print.hh>

				static inline int8_t hex_to_int(unsigned char c) {

				@@ -42,7 +43,7 @@ bytes from_hex(sstring_view s) {

				        auto half_byte1 = hex_to_int(s[i * 2]);

				        auto half_byte2 = hex_to_int(s[i * 2 + 1]);

				        if (half_byte1 == -1 || half_byte2 == -1) {

				            throw std::invalid_argument(format("Non-hex characters in {}", s));

				            throw std::invalid_argument(fmt::format("Non-hex characters in {}", s));

				        }

				        out[i] = (half_byte1 << 4) | half_byte2;

				    }

									
										5

bytes.hh
									
												View File
												
				@@ -16,13 +16,10 @@

				#include <iosfwd>

				#include <functional>

				#include <compare>

				#include "bytes_fwd.hh"

				#include "utils/mutable_view.hh"

				#include "utils/simple_hashers.hh"

				using bytes = basic_sstring<int8_t, uint32_t, 31, false>;

				using bytes_view = std::basic_string_view<int8_t>;

				using bytes_mutable_view = basic_mutable_view<bytes_view::value_type>;

				using bytes_opt = std::optional<bytes>;

				using sstring_view = std::string_view;

				inline bytes to_bytes(bytes&& b) {

									
										20

bytes_fwd.hh
									
										Normal file
									
												View File
												
				@@ -0,0 +1,20 @@

				/*

				 * Copyright (C) 2024-present ScyllaDB

				 */

				/*

				 * SPDX-License-Identifier: AGPL-3.0-or-later

				 */

				#pragma once

				#include <seastar/core/sstring.hh>

				#include "utils/mutable_view.hh"

				using namespace seastar;

				using bytes = basic_sstring<int8_t, uint32_t, 31, false>;

				using bytes_view = std::basic_string_view<int8_t>;

				using bytes_mutable_view = basic_mutable_view<bytes_view::value_type>;

				using bytes_opt = std::optional<bytes>;

									
										93

bytes_ostream.hh
									
												View File
												
				@@ -8,14 +8,58 @@

				#pragma once

				#include <boost/range/iterator_range.hpp>

				#include "bytes.hh"

				#include "utils/assert.hh"

				#include "utils/managed_bytes.hh"

				#include <seastar/core/simple-stream.hh>

				#include <seastar/core/loop.hh>

				#include <bit>

				#include <concepts>

				#include <ranges>

				class bytes_ostream_fragment_iterator {

				public:

				    using iterator_category = std::input_iterator_tag;

				    using iterator_concept = std::input_iterator_tag;

				    using value_type = bytes_view;

				    using difference_type = std::ptrdiff_t;

				    using pointer = bytes_view*;

				    using reference = bytes_view&;

				public:

				    using chunk = multi_chunk_blob_storage;

				    struct implementation {

				        chunk* current_chunk;

				    };

				private:

				    chunk* _current = nullptr;

				public:

				    bytes_ostream_fragment_iterator() = default;

				    bytes_ostream_fragment_iterator(chunk* current) : _current(current) {}

				    bytes_ostream_fragment_iterator(const bytes_ostream_fragment_iterator&) = default;

				    bytes_ostream_fragment_iterator& operator=(const bytes_ostream_fragment_iterator&) = default;

				    bytes_view operator*() const {

				        return { _current->data, _current->frag_size };

				    }

				    bytes_view operator->() const {

				        return *(*this);

				    }

				    bytes_ostream_fragment_iterator& operator++() {

				        _current = _current->next;

				        return *this;

				    }

				    bytes_ostream_fragment_iterator operator++(int) {

				        bytes_ostream_fragment_iterator tmp(*this);

				        ++(*this);

				        return tmp;

				    }

				    bool operator==(const bytes_ostream_fragment_iterator&) const = default;

				    implementation extract_implementation() const {

				        return implementation {

				            .current_chunk = _current,

				        };

				    }

				};

				/**

				 * Utility for writing data into a buffer when its final size is not known up front.

				@@ -45,46 +89,7 @@ private:

				    size_type _size;

				    size_type _initial_chunk_size = default_chunk_size;

				public:

				    class fragment_iterator {

				    public:

				        using iterator_category = std::input_iterator_tag;

				        using value_type = bytes_view;

				        using difference_type = std::ptrdiff_t;

				        using pointer = bytes_view*;

				        using reference = bytes_view&;

				        struct implementation {

				            chunk* current_chunk;

				        };

				    private:

				        chunk* _current = nullptr;

				    public:

				        fragment_iterator() = default;

				        fragment_iterator(chunk* current) : _current(current) {}

				        fragment_iterator(const fragment_iterator&) = default;

				        fragment_iterator& operator=(const fragment_iterator&) = default;

				        bytes_view operator*() const {

				            return { _current->data, _current->frag_size };

				        }

				        bytes_view operator->() const {

				            return *(*this);

				        }

				        fragment_iterator& operator++() {

				            _current = _current->next;

				            return *this;

				        }

				        fragment_iterator operator++(int) {

				            fragment_iterator tmp(*this);

				            ++(*this);

				            return tmp;

				        }

				        bool operator==(const fragment_iterator&) const = default;

				        implementation extract_implementation() const {

				            return implementation {

				                .current_chunk = _current,

				            };

				        }

				    };

				    using fragment_iterator = bytes_ostream_fragment_iterator;

				    using const_iterator = fragment_iterator;

				    class output_iterator {

				@@ -269,7 +274,7 @@ public:

				    // Call only when is_linearized()

				    bytes_view view() const {

				        assert(is_linearized());

				        SCYLLA_ASSERT(is_linearized());

				        if (!_current) {

				            return bytes_view();

				        }

				@@ -357,7 +362,7 @@ public:

				    output_iterator write_begin() { return output_iterator(*this); }

				    boost::iterator_range<fragment_iterator> fragments() const {

				    std::ranges::subrange<fragment_iterator> fragments() const {

				        return { begin(), end() };

				    }

									
										8

cache_mutation_reader.hh
									
												View File
												
				@@ -8,6 +8,7 @@

				#pragma once

				#include "utils/assert.hh"

				#include <vector>

				#include "row_cache.hh"

				#include "mutation/mutation_fragment.hh"

				@@ -283,7 +284,7 @@ future<> cache_mutation_reader::process_static_row() {

				        return ensure_underlying().then([this] {

				            return (*_underlying)().then([this] (mutation_fragment_v2_opt&& sr) {

				                if (sr) {

				                    assert(sr->is_static_row());

				                    SCYLLA_ASSERT(sr->is_static_row());

				                    maybe_add_to_cache(sr->as_static_row());

				                    push_mutation_fragment(std::move(*sr));

				                }

				@@ -382,7 +383,7 @@ future<> cache_mutation_reader::do_fill_buffer() {

				    if (_state == state::reading_from_underlying) {

				        return read_from_underlying();

				    }

				    // assert(_state == state::reading_from_cache)

				    // SCYLLA_ASSERT(_state == state::reading_from_cache)

				    return _lsa_manager.run_in_read_section([this] {

				        auto next_valid = _next_row.iterators_valid();

				        clogger.trace("csm {}: reading_from_cache, range=[{}, {}), next={}, valid={}, rt={}", fmt::ptr(this), _lower_bound,

				@@ -794,7 +795,6 @@ void cache_mutation_reader::copy_from_cache_to_buffer() {

				            };

				            if (row_tomb_expired(t) || is_row_dead(row)) {

				                can_gc_fn always_gc = [&](tombstone) { return true; };

				                const schema& row_schema = _next_row.latest_row_schema();

				                _read_context.cache()._tracker.on_row_compacted();

				@@ -990,7 +990,7 @@ void cache_mutation_reader::offer_from_underlying(mutation_fragment_v2&& mf) {

				        maybe_add_to_cache(mf.as_clustering_row());

				        add_clustering_row_to_buffer(std::move(mf));

				    } else {

				        assert(mf.is_range_tombstone_change());

				        SCYLLA_ASSERT(mf.is_range_tombstone_change());

				        auto& chg = mf.as_range_tombstone_change();

				        if (maybe_add_to_cache(chg)) {

				            add_to_buffer(std::move(mf).as_range_tombstone_change());

									
										6

cartesian_product.hh
									
												View File
												
				@@ -76,7 +76,7 @@ public:

				};

				template<typename T>

				static inline

				inline

				size_t cartesian_product_size(const std::vector<std::vector<T>>& vec_of_vecs) {

				    size_t r = 1;

				    for (auto&& vec : vec_of_vecs) {

				@@ -86,7 +86,7 @@ size_t cartesian_product_size(const std::vector<std::vector<T>>& vec_of_vecs) {

				}

				template<typename T>

				static inline

				inline

				bool cartesian_product_is_empty(const std::vector<std::vector<T>>& vec_of_vecs) {

				    for (auto&& vec : vec_of_vecs) {

				        if (vec.empty()) {

				@@ -97,7 +97,7 @@ bool cartesian_product_is_empty(const std::vector<std::vector<T>>& vec_of_vecs)

				}

				template<typename T>

				static inline

				inline

				cartesian_product<T> make_cartesian_product(const std::vector<std::vector<T>>& vec_of_vecs) {

				    return cartesian_product<T>(vec_of_vecs);

				}

									
										4

cdc/cdc_extension.hh
									
												View File
												
				@@ -11,7 +11,7 @@

				#include <seastar/core/sstring.hh>

				#include "bytes.hh"

				#include "bytes_fwd.hh"

				#include "cdc/cdc_options.hh"

				#include "schema/schema.hh"

				#include "serializer_impl.hh"

				@@ -34,7 +34,7 @@ public:

				        return ser::serialize_to_buffer<bytes>(_cdc_options.to_map());

				    }

				    static std::map<sstring, sstring> deserialize(const bytes_view& buffer) {

				        return ser::deserialize_from_buffer(buffer, boost::type<std::map<sstring, sstring>>());

				        return ser::deserialize_from_buffer(buffer, std::type_identity<std::map<sstring, sstring>>());

				    }

				    const options& get_options() const {

				        return _cdc_options;

									
										2

cdc/cdc_partitioner.cc
									
												View File
												
				@@ -23,7 +23,7 @@ const sstring cdc_partitioner::name() const {

				}

				static dht::token to_token(int64_t value) {

				    return dht::token(dht::token::kind::key, value);

				    return dht::token(value);

				}

				static dht::token to_token(bytes_view key) {

Compare commits

1830 Commits next-6.1 ... dani-tweig

209 .clang-format Normal file Unescape Escape View File

31 .github/CODEOWNERS vendored Unescape Escape View File

15 .github/ISSUE_TEMPLATE.md vendored Unescape Escape View File

86 .github/ISSUE_TEMPLATE/bug_report.yml vendored Normal file Unescape Escape View File

9 .github/dependabot.yml vendored Normal file Unescape Escape View File

58 .github/mergify.yml vendored Unescape Escape View File

181 .github/scripts/auto-backport.py vendored Executable file Unescape Escape View File

23 .github/scripts/label_promoted_commits.py vendored Unescape Escape View File

51 .github/workflows/add-label-when-promoted.yaml vendored Unescape Escape View File

9 .github/workflows/backport-pr-fixes-validation.yaml vendored Unescape Escape View File

8 .github/workflows/build-scylla.yaml vendored Unescape Escape View File

3 .github/workflows/clang-nightly.yaml vendored Unescape Escape View File

7 .github/workflows/clang-tidy.yaml vendored Unescape Escape View File

45 .github/workflows/conflict_reminder.yaml vendored Normal file Unescape Escape View File

3 .github/workflows/docs-pr.yaml vendored Unescape Escape View File

4 .github/workflows/iwyu.yaml vendored Unescape Escape View File

1 .github/workflows/reproducible-build.yaml vendored Unescape Escape View File

5 .gitignore vendored Unescape Escape View File

3 .gitmodules vendored Unescape Escape View File

92 CMakeLists.txt Unescape Escape View File

19 HACKING.md Unescape Escape View File

10 README.md Unescape Escape View File

4 SCYLLA-VERSION-GEN Unescape Escape View File

25 alternator/auth.cc Unescape Escape View File

14 alternator/conditions.cc Unescape Escape View File

6 alternator/controller.cc Unescape Escape View File

502 alternator/executor.cc View File

15 alternator/executor.hh Unescape Escape View File

19 alternator/expressions.cc Unescape Escape View File

2 alternator/rmw_operation.hh Unescape Escape View File

37 alternator/serialization.cc Unescape Escape View File

61 alternator/server.cc Unescape Escape View File

9 alternator/server.hh Unescape Escape View File

6 alternator/stats.cc Unescape Escape View File

4 alternator/stats.hh Unescape Escape View File

32 alternator/streams.cc Unescape Escape View File

114 alternator/ttl.cc Unescape Escape View File

31 api/CMakeLists.txt Unescape Escape View File

8 api/api-doc/column_family.json Unescape Escape View File

26 api/api-doc/cql_server_test.json Normal file Unescape Escape View File

32 api/api-doc/raft.json Unescape Escape View File

156 api/api-doc/storage_service.json Unescape Escape View File

15 api/api-doc/system.json Unescape Escape View File

55 api/api-doc/task_manager.json Unescape Escape View File

54 api/api.cc Unescape Escape View File

2 api/api.hh Unescape Escape View File

13 api/api_init.hh Unescape Escape View File

45 api/cache_service.cc Unescape Escape View File

1 api/cache_service.hh Unescape Escape View File

3 api/collectd.cc Unescape Escape View File

35 api/column_family.cc Unescape Escape View File

2 api/column_family.hh Unescape Escape View File

43 api/commitlog.cc Unescape Escape View File

7 api/commitlog.hh Unescape Escape View File

62 api/compaction_manager.cc Unescape Escape View File

6 api/compaction_manager.hh Unescape Escape View File

12 api/config.cc Unescape Escape View File

69 api/cql_server_test.cc Normal file Unescape Escape View File

29 api/cql_server_test.hh Normal file Unescape Escape View File

2 api/lsa.cc Unescape Escape View File

48 api/raft.cc Unescape Escape View File

113 api/storage_service.cc Unescape Escape View File

1 api/storage_service.hh Unescape Escape View File

15 api/system.cc Unescape Escape View File

5 api/system.hh Unescape Escape View File

264 api/task_manager.cc Unescape Escape View File

39 api/task_manager_test.cc Unescape Escape View File

1 api/tasks.cc Unescape Escape View File

5 api/token_metadata.cc Unescape Escape View File

57 auth/authentication_options.hh Unescape Escape View File

18 auth/authenticator.hh Unescape Escape View File

6 auth/certificate_authenticator.cc Unescape Escape View File

1 auth/certificate_authenticator.hh Unescape Escape View File

14 auth/common.cc Unescape Escape View File

6 auth/common.hh Unescape Escape View File

28 auth/default_authorizer.cc Unescape Escape View File

15 auth/maintenance_socket_role_manager.cc Unescape Escape View File

6 auth/maintenance_socket_role_manager.hh Unescape Escape View File

1830 Commits

next-6.1 ... dani-tweig

209

.clang-format Normal file

View File

31

.github/CODEOWNERS vendored

View File

15

.github/ISSUE_TEMPLATE.md vendored

View File

86

.github/ISSUE_TEMPLATE/bug_report.yml vendored Normal file

View File

9

.github/dependabot.yml vendored Normal file

View File

58

.github/mergify.yml vendored

View File

181

.github/scripts/auto-backport.py vendored Executable file

View File

23

.github/scripts/label_promoted_commits.py vendored

View File

51

.github/workflows/add-label-when-promoted.yaml vendored

View File

9

.github/workflows/backport-pr-fixes-validation.yaml vendored

View File

8

.github/workflows/build-scylla.yaml vendored

View File

3

.github/workflows/clang-nightly.yaml vendored

View File

7

.github/workflows/clang-tidy.yaml vendored

View File

45

.github/workflows/conflict_reminder.yaml vendored Normal file

View File

3

.github/workflows/docs-pr.yaml vendored

View File

4

.github/workflows/iwyu.yaml vendored

View File

1

.github/workflows/reproducible-build.yaml vendored

View File

5

.gitignore vendored

View File

3

.gitmodules vendored

View File

92

CMakeLists.txt

View File

19

HACKING.md

View File

10

README.md

View File

4

SCYLLA-VERSION-GEN

View File

25

alternator/auth.cc

View File

14

alternator/conditions.cc

View File

6

alternator/controller.cc

View File

502

alternator/executor.cc

View File

15

alternator/executor.hh

View File

19

alternator/expressions.cc

View File

2

alternator/rmw_operation.hh

View File

37

alternator/serialization.cc

View File

61

alternator/server.cc

View File

9

alternator/server.hh

View File

6

alternator/stats.cc

View File

4

alternator/stats.hh

View File

32

alternator/streams.cc

View File

114

alternator/ttl.cc

View File

31

api/CMakeLists.txt

View File

8

api/api-doc/column_family.json

View File

26

api/api-doc/cql_server_test.json Normal file

View File

32

api/api-doc/raft.json

View File

156

api/api-doc/storage_service.json

View File

15

api/api-doc/system.json

View File

55

api/api-doc/task_manager.json

View File

54

api/api.cc

View File

2

api/api.hh

View File

13

api/api_init.hh

View File

45

api/cache_service.cc

View File

1

api/cache_service.hh

View File

3

api/collectd.cc

View File

35

api/column_family.cc

View File

2

api/column_family.hh

View File

43

api/commitlog.cc

View File

7

api/commitlog.hh

View File

62

api/compaction_manager.cc

View File

6

api/compaction_manager.hh

View File

12

api/config.cc

View File

69

api/cql_server_test.cc Normal file

View File

29

api/cql_server_test.hh Normal file

View File

2

api/lsa.cc

View File

48

api/raft.cc

View File

113

api/storage_service.cc

View File

1

api/storage_service.hh

View File

15

api/system.cc

View File

5

api/system.hh

View File

264

api/task_manager.cc

View File

39

api/task_manager_test.cc

View File

1

api/tasks.cc

View File

5

api/token_metadata.cc

View File

57

auth/authentication_options.hh

View File

18

auth/authenticator.hh

View File

6

auth/certificate_authenticator.cc

View File

1

auth/certificate_authenticator.hh

View File

14

auth/common.cc

View File

6

auth/common.hh

View File

28

auth/default_authorizer.cc

View File

15

auth/maintenance_socket_role_manager.cc

View File

6

auth/maintenance_socket_role_manager.hh

View File

100

auth/password_authenticator.cc

View File