scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-05-12 19:02:12 +00:00

Author	SHA1	Message	Date
Kefu Chai	f5b29331a2	build: populate --enable-dist --disable-dist to CMake before this change, the "dist" targets are always enabled in the CMake-based building system. but the build rules generated by `configure.py` does respect `--enable-dist` and `--disable-dist` command line options, and enable/distable the dist targets respectively. in this change, we - add an CMake option named "Scylla_DIST". the "dist" subdirectory in CMake only if this option is ON. - pouplate the `--enable-dist` and `--disable-dist` option down to cmake by setting the `Scylla_DIST` option, when creating the build system using CMake. this enables the CMake-based build system to be functionality wise more closer to the legacy building system. Refs scylladb/scylladb#2717 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21253	2024-10-27 21:57:46 +02:00
Avi Kivity	3124711fc4	Merge 'Report rows_merged in compaction_history rest api and nodetool' from Łukasz Paszkowski Currently, running the `nodetool compactionhistory` command or using the rest api `curl -X GET --header "Accept: application/json" "http://localhost:10000/compaction_manager/compaction_history"` return compaction history without the `row_merged` field. The series computes rows merged during compaction and provides this information to users via both the nodetool command and the rest api. The `rows_merged` field contains information on merged clustering keys across multiple sstable files. For instance, compacting two sstables of a table consisting of 7 rows where two rows are part of the both sstables, the output would have the following format: {1: 5, 2: 2}. No backport is required. It extends the existing compaction history output. Fixes https://github.com/scylladb/scylladb/issues/666 Closes scylladb/scylladb#20481 * github.com:scylladb/scylladb: test/rest_api: Add tests for compactionhistory nodetool: Add rows merged stats into compactionhistory output compaction: Update compaction history with collected histogram compaction: Remove const qualifier from methods creating sstable readers sstable_set: Add optional statistics to make_local_shard_sstable_reader make_combined_reader: Add optional parameter, combined_reader_statistics reader_selector: Extend with maximum reader count mutation_fragment_merger: Create histogram while consuming mutation fragment batches	2024-10-27 21:26:11 +02:00
Kefu Chai	a9e18f70b0	Revert submodule change in `6ead5a4696` in `6ead5a46`, we included submodule changes in cqlsh and java by accident. this was not intended. and this broke the artifacts-rocky8-test. in this change, both changes in the submodule are reverted. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21236	2024-10-23 19:49:20 +03:00
Łukasz Paszkowski	8188a71787	nodetool: Add rows merged stats into compactionhistory output Incorporate rows merged statistics into the output of the compactionhistory command. Depending on the requested format type, the output has different form. For instance, compacting two sstables of a table consisting of 7 rows where two rows are part of the both sstables, the output would have the following format: text: {1: 5, 2: 2} json: [{"key":1,"value":5},{"key":2,"value":1}]} yaml: - key: 1 value: 5 - key: 2 value: 1	2024-10-22 08:15:02 +02:00
Kefu Chai	6ead5a4696	treewide: move log.hh into utils/log.hh the log.hh under the root of the tree was created keep the backward compatibility when seastar was extracted into a separate library. so log.hh should belong to `utils` directory, as it is based solely on seastar, and can be used all subsystems. in this change, we move log.hh into utils/log.hh to that it is more modularized. and this also improves the readability, when one see `#include "utils/log.hh"`, it is obvious that this source file needs the logging system, instead of its own log facility -- please note, we do have two other `log.hh` in the tree. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-22 06:54:46 +03:00
Kefu Chai	5cd619a60c	treewide: s/boost::adaptors::map_keys/std::views::keys/ now that we are allowed to use C++23. we now have the luxury of using `std::views::keys`. in this change, we: - replace `boost::adaptors::map_keys` with `std::views::keys` - update affected code to work with `std::views::keys` to reduce the dependency to boost for better maintainability, and leverage standard library features for better long-term support. this change is part of our ongoing effort to modernize our codebase and reduce external dependencies where possible. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21198	2024-10-21 12:47:52 +03:00
Avi Kivity	c3be2489ce	treewide: drop includes of <boost/range/adaptors.hpp> This includes way too much, including <boost/regex.hpp>, which is huge. Drop includes of adaptors.hpp and replace by what is needed. Closes scylladb/scylladb#21187	2024-10-20 17:17:11 +03:00
Kefu Chai	5ef0cbb693	tools/scylla-nodetool: s/vm.count()/vm.contains()/ this change is created in the same spirit of `0104c7d3`, which used `std::map::contains()` in the place of `std::map::count()` when checking for the existence of a paramter with given name for better readability. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21158	2024-10-17 13:41:15 +03:00
Kefu Chai	d7f315ef63	tool/scylla-nodetool: check for positional argument passed to "restore" before this change, if no positional arguments are passed to "restore" subcommand, the tool fails with following error message: ``` error running operation: boost::wrapexcept<boost::bad_any_cast> (boost::bad_any_cast: failed conversion using boost::any_cast) ``` this is difficult to digest. after this change, if no sstables are specified: ``` error processing arguments: missing required parameter: sstables ``` this is slightly better from user experience's perspective. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21136	2024-10-17 13:41:15 +03:00
Avi Kivity	f799234c82	Update tools/java submodule (deprecation notice) * tools/java b2d025fd6b...807e991de7 (1): > README.md: add deprecation notice for java tools	2024-10-16 17:09:48 +03:00
Sergey Zolotukhin	ae23d42889	tools: Add build_info header with functions providing build type information A new header provides `constexpr` functions to retrieve build type information: `get_build_type()`, `is_release_build()`, and `is_debug_build()`. These functions are useful when adding changes that should be enabled at compile time only for specific build types.	2024-10-11 09:38:24 +02:00
Benny Halevy	3a12ad96c7	sstables: scylla_metadata: add sstable identifier Keep a copy of the sstable uuid generation in a new scylla_metadata sstable_identifier attribute. If the SSTable happens to have a numerical generation just create a new time-uuid and log a message about that. Dump this new attribute in scylla sstable dump tool. And add a unit test to verify that the written (and then loaded) sstable identifier matches the sstable's generation. The motivatrion for this change stems from backup deduplication. In essence, an sstable may already have been backed up in a previous snapshot, and we don't want to abck it up again if it's already present on external storage. Today this is based on rclone that compares files checksums, but once scylla will backup the sstables using the native object-storage stack (#19890), we would like to use the sstable globally-unique identifier for deduplication. Although the uuid-generation is encoded in the sstable path, the latter may change, e.g. due to intra-node migration, so keep a copy of the original unique identifier in scylla-metadata, and that attribute would survive file-based or intra-node migrations. Fixes scylladb/scylladb#20459 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes scylladb/scylladb#21002	2024-10-10 08:52:46 +03:00
Yuao Ma	1cc7821d12	tools: fix typos in the code This patch corrects a minor typo without any functional changes. Signed-off-by: Yuao Ma <c8ef@outlook.com> Closes scylladb/scylladb#20975	2024-10-09 08:18:36 +03:00
Pavel Emelyanov	6b480589fe	Merge 'treewide: accept list of sstables in "restore" API ' from Kefu Chai before this change, we enumerate the sstables tracked by the system.sstables table, and restore them when serving requests to "storage_service/restore" API. this works fine with "storage_service/backup" API. but this "restore" API cannot be used as a drop-in replacement of the rclone based API currently used by scylla-manager. in order to fill the gap, in this change: * add the "prefix" parameter for specifying the shared prefix of sstables * add the "sstables" parameter for specifying the list of TOC components of sstables * remove the "snapshot" parameter, as we don't encode the prefix on scylla's end anymore. * make the "table" parameter mandatory. Fixes https://github.com/scylladb/scylladb/issues/20461 ---- this change is a part of the efforts to bring the native backup/restore to scylla, no need to backprt. Closes scylladb/scylladb#20685 * github.com:scylladb/scylladb: treewide: accept list of sstables in "restore" API sstable: pass get_storage_option to sstable_directory::load_sstable() test/nodetool: add body parameter to `expected_request` tools/scylla-nodetool: enable nodetool to write HTTP body	2024-10-04 12:38:08 +03:00
Botond Dénes	07094c3e44	Merge 'replica: Fix tombstone GC during tablet split preparation' from Raphael "Raph" Carvalho During split prepare phase, there will be more than 1 compaction group with overlapping token range for a given replica. Assume tablet 1 has sstable A containing deleted data, and sstable B containing a tombstone that shadows data in A. Then split starts: 1) sstable B is split first, and moved from main (unsplit) group to a split-ready group 2) now compaction runs in split-ready group before sstable A is split tombstone GC logic today only looks at underlying group, so compaction is step 2 will discard the deleted data in A, since it belongs to another group (the unsplit one), and so the tombstone can be purged incorrectly. To fix it, compaction will now work with all uncompacting sstables that belong to the same replica, since tombstone GC requires all sstables that possibly contain shadowed data to be available for correct decision to be made. Fixes https://github.com/scylladb/scylladb/issues/20044. Branches 6.0, 6.1 and 6.2 are vulnerable, so backport is needed. Closes scylladb/scylladb#20939 * github.com:scylladb/scylladb: replica: Fix tombstone GC during tablet split preparation service: Improve error handling for split	2024-10-04 10:29:42 +03:00
Raphael S. Carvalho	93815e0649	replica: Fix tombstone GC during tablet split preparation During split prepare phase, there will be more than 1 compaction group with overlapping token range for a given replica. Assume tablet 1 has sstable A containing deleted data, and sstable B containing a tombstone that shadows data in A. Then split starts: 1) sstable B is split first, and moved from main (unsplit) group to a split-ready group 2) now compaction runs in split-ready group before sstable A is split tombstone GC logic today only looks at underlying group, so compaction is step 2 will discard the deleted data in A, since it belongs to another group (the unsplit one), and so the tombstone can be purged incorrectly. To fix it, compaction will now work with all uncompacting sstables that belong to the same replica, since tombstone GC requires all sstables that possibly contain shadowed data to be available for correct decision to be made. Fixes #20044. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-10-02 11:26:13 -03:00
Sergey Zolotukhin	9c692438e9	nodetool: Add IP address usage warning for 'ignore-dead-nodes'. Since we are deprecating the use of IP addresses, a warning message will be printed if 'nodetool removenode --ignore-dead-nodes' is used with IP addresses.	2024-10-02 11:56:59 +02:00
Kefu Chai	787ea4b1d4	treewide: accept list of sstables in "restore" API before this change, we enumerate the sstables tracked by the system.sstables table, and restore them when serving requests to "storage_service/restore" API. this works fine with "storage_service/backup" API. but this "restore" API cannot be used as a drop-in replacement of the rclone based API currently used by scylla-manager. in order to fill the gap, in this change: * add the "prefix" parameter for specifying the shared prefix of sstables * add the "sstables" parameter for specifying the list of TOC components of sstables * remove the "snapshot" parameter, as we don't encode the prefix on scylla's end anymore. * make the "table" parameter mandatory. Fixes scylladb/scylladb#20461 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-01 23:24:56 +08:00
Kefu Chai	3c19cc9aec	tools/scylla-nodetool: enable nodetool to write HTTP body before this change, we always send the parameters with query strings, but we will add an API ("storage_service/restore") which accepts its parameters in HTTP body as well. in this change, we add an optional parameter to `do_request()` and `post()`, so that we can send HTTP body when using "POST" method in nodetool implementation. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-10-01 23:24:56 +08:00
Avi Kivity	f5628be597	Update tools/java submodule * tools/java 5b0e274f12...b2d025fd6b (1): > build.xml: update scylla-tools license	2024-10-01 12:48:45 +03:00
Yaron Kaikov	d164fd45bc	install-dependencies.sh: update node_exporter to 1.8.2 Update node_exporter to 1.8.2 Fixes: #18493 Closes scylladb/scylladb#20254 [avi: regenerate frozen toolchain, with new clang in https://devpkg.scylladb.com/clang/clang-18.1.8-Fedora-40-aarch64.tar.gz https://devpkg.scylladb.com/clang/clang-18.1.8-Fedora-40-x86_64.tar.gz new clang regenerated due to new packaging format (`f6fe4d9e73`) and some other minor changes.]	2024-09-25 18:42:25 +03:00
Pavel Emelyanov	ae76481444	Merge 'treewide: add "table" parameter to "backup" API ' from Kefu Chai with this parameter, "backup" API can backup the given table, this enables it to be a drop-in replacement of existing rclone API used by scylla manager. Fixes https://github.com/scylladb/scylladb/issues/20636 --- this change is a part of the efforts to bring the native backup/restore to scylla, no need to backprt. Closes scylladb/scylladb#20661 * github.com:scylladb/scylladb: backup_task: fix the indent treewide: add "table" parameter to "backup" API	2024-09-25 10:53:38 +03:00
Takuya ASADA	f6fe4d9e73	toolchain: fix broken INSTALL_FROM mode We found that --clang-build-mode INSTALL_FROM tries to rebuild clang even we use an archive of prebuilt image. Seems like it is because ninja detected changes on standard library headers, which updated when we build new frozen toolchain container image. To avoid such unnecessary rebuild, we should stop archive whole clang build directory, we should archive install image instead. To do so, we can use "DESTDIR=<sysroot dir> ninja install-distribution-stripped", and archive sysroot dir as clang archive. Fixes #20421 Closes scylladb/scylladb#20422	2024-09-25 10:48:56 +03:00
Kefu Chai	d663b6c13b	treewide: add "table" parameter to "backup" API with this parameter, "backup" API can backup the given table, this enables it to be a drop-in replacement of existing rclone API used by scylla manager. in this change: * api/storage_service: add "table" parameter to "backup" API. * snapshot_ctl: compose the full path of the snapshot directory in `snapshot_ctl::start_backup`. since we have all the information for composing the snapshot directory, and what the `backup_task_impl` class is interested is but the snapshot directory, we just pass the path to it instead the individual components of the directory. * backup_task_impl: instead of scan the whole keyspace recursively, only scan the specified snapshot directory. Fixes scylladb/scylladb#20636 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-25 09:11:26 +08:00
Avi Kivity	d16ea0afd6	Merge 'cql3: Extend DESC SCHEMA by auth and service levels' from Dawid Mędrek Auth has been managed via Raft since Scylla 6.0. Restoring data following the usual procedure (1) is error-prone and so a safer method must have been designed and implemented. That's what happens in this PR. We want to extend `DESC SCHEMA` by auth and service levels to provide a safe way to backup and restore those two components. To realize that, we change the meaning of `DESC SCHEMA WITH INTERNALS` and add a new "tier": `DESC SCHEMA WITH INTERNALS AND PASSWORDS`. * `DESC SCHEMA` -- no change, i.e. the statement describes the current schema items such as keyspaces, tables, views, UDTs, etc. * `DESC SCHEMA WITH INTERNALS` -- does the same as the previous tier and also describes auth and service levels. No information about passwords is returned. * `DESC SCHEMA WITH INTERNALS AND PASSWORDS` -- does the same as the previous tier and also includes information about the salted hashes corresponding to the passwords of roles. To restore existing roles, we extend the `CREATE ROLE` statement by allowing to use the option `WITH SALTED HASH = '[...]'`. --- Implementation strategy: * Add missing things/adjust existing ones that will be used later. * Implement creating a role with salted hash. * Add tests for creating a role with salted hash. * Prepare for implementing describe functionality of auth and service levels. * Implement describe functionality for elements of auth and service levels. * Extend the grammar. * Add tests for describe auth and service levels. * Add/update documentation. --- (1): https://opensource.docs.scylladb.com/stable/operating-scylla/procedures/backup-restore/restore.html In case the link stops working, restoring a schema was realised by managing raw files on disk. Fixes scylladb/scylladb#18750 Fixes scylladb/scylladb#18751 Fixes scylladb/scylladb#20711 Closes scylladb/scylladb#20168 * github.com:scylladb/scylladb: docs: Update user documentation for backup and restore docs/dev: Add documentation for DESC SCHEMA test: Add tests for describing auth and service levels cql3/functions/user_function: Remove newline character before and after UDF body cql3: Implement DESCRIBE SCHEMA WITH INTERNALS AND PASSWORDS auth: Implement describing auth auth/authenticator: Add member functions for querying password hash service/qos/service_level_controller: Describe service levels data_dictionary: Remove keyspace_element.hh treewide: Start using new overloads of describe treewide: Fix indentation in describe functions treewide: Return create statement optionally in describe functions treewide: Add new describe overloads to implementations of data_dictionary::keyspace_element treewide: Start using schema::ks_name() instead of schema::keyspace_name() cql3: Refactor `description` cql3: Move description to dedicated files test: Add tests for `CREATE ROLE WITH SALTED HASH` cql3/statements: Restrict CREATE ROLE WITH SALTED HASH auth: Allow for creating roles with SALTED HASH types: Introduce a function `cql3_type_name_without_frozen()` cql3/util: Accept std::string_view rather than const sstring&	2024-09-24 21:44:32 +03:00
Yaniv Michael Kaul	26f2cbdfe2	optimized_clang.sh: compile with -march Add for both x86_64 compilation flags for clang, to get it compile with newer arch x86_64-v3 for x86 and ARM 8.2 level for aarch64. Tested to compile fine with both clang 18.1.6 and 18.1.8. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Closes scylladb/scylladb#20682	2024-09-23 17:40:20 +03:00
Yaniv Michael Kaul	85c0bb7ff4	optimized_clang.sh: add missing symbolic links to clang (for ccache) The removal of clang removes the symblic links ccache uses to mask itself as clang/clang++ Manually add them back, so ccache can work. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com> Fixes: https://github.com/scylladb/scylladb/issues/20490 Closes scylladb/scylladb#20491	2024-09-23 15:55:22 +03:00
Kefu Chai	40f2d4c988	build: cmake: drop scylla-jmx from the build in `3cd2a61736`, we dropped scylla-jmx from the build. but didn't update the CMake building system accordingly, this broke the CMake build, as the dependencies pointing to jmx cannot be found or fulfilled. in this change, we remove all references to jmx in the CMake build. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com> Closes scylladb/scylladb#20736	2024-09-23 14:20:42 +03:00
Avi Kivity	3d781c4fc8	Update frozen toolchain * tools/java e505a6d3bb...5b0e274f12 (1): > Merge 'build.xml: install and use java-11 when building' from Kefu Chai Updates to clang 18.1.8 + LLVM patch to match Fedora 40. New optimized clang build generated and stored in https://devpkg.scylladb.com/clang/clang-18.1.8-x86_64.tar.gz https://devpkg.scylladb.com/clang/clang-18.1.8-aarch64.tar.gz Due to the loss of the jmx submodule, we no longer install java-11-openjdk. We add it in install-dependencies.sh here to compensate, pending a better solution. tools/java submodule updated to remove build failure where Java 8 was selected instead of Java 11. The scylla_gdb test suite was disabled due to a regression in gdb 15, which is brought in by the toolchain update [1]. [1] https://github.com/scylladb/scylladb/issues/20741.	2024-09-21 20:07:28 +03:00
Botond Dénes	488a372fdc	tool/scylla-nodetool: status: reorder endpoint calls to match old nodetool Old nodetool requested `/storage_service/tokens_endpoing` first, then `/storage_service/host_id`, while the native nodetool did it in reverse order. Most of the time this is inconsequential but there is an edge case when a node's IP address is changed. This reversing of the order results in unexpected behavior for tests, causing noise via flaky tests. Match the order of the old nodetool so that the native nodetool exhibits the behavior expected by tests (and users too probably). Fixes: scylladb/scylladb#18693 Closes scylladb/scylladb#20615	2024-09-20 15:07:16 +02:00
Dawid Mędrek	39cf106151	treewide: Start using schema::ks_name() instead of schema::keyspace_name() We're going to remove the interface `data_dictionary::keyspace_element`. As `schema::keyspace_name()` is an implementation of one of the methods specified by that interface, we replace its uses by `schema::ks_name()`. `schema::keyspace_name()` was an alias for it, so no semantic change has occured.	2024-09-20 14:24:53 +02:00
Avi Kivity	61d19e4464	Update tools/java submodule * tools/java 0b4accdd5e...e505a6d3bb (1): > [C-S] Make it use DCAwareRoundRobinPolicy unless rack is provided	2024-09-20 14:49:21 +03:00
Ernest Zaslavsky	924325fd25	treewide: add "prefix" parameter to backup API Allow the caller to pass the prefix when performing backup and restore Fixes scylladb/scylladb#20335 Closes scylladb/scylladb#20413	2024-09-18 08:25:00 +03:00
Andrei Chekun	bbb6c3c2ff	test.py: Add resource consumption metrics This PR adds the possibility to gather resource consumption metrics. The collected metrics can be used to compare performance before and after specific changes aimed at increasing performance. Currently, this functionality works only in manual mode, and this is just raw data. Later on, these metrics can be used in Jupyter notebook to analyze and visualize how the resources are used and can provide the insight on how to improve it. This PR is a first insight after gathering these metrics. Add the possibility to gather resource consumption for the test.py execution. SQLite DB will be created with different performance metrics that will allow comparing the resource consumption between changes. The DB will be in the tmp directory that by default set to testlog. Across the runs, the DB will not be deleted, so each new run will just add information to the existing DB. Parameter --get-metrics was added to switch on or off the metrics gathering. By default, it's switched on. Closes: scylladb/qa-tasks#1666 Closes: scylladb/qa-tasks#1707 Closes scylladb/scylladb#19881	2024-09-17 15:22:34 +03:00
Botond Dénes	6250ff18eb	Merge 'sstable: s/crawling_sstable_mutation_reader/sstable_full_scan_reader' from Kefu Chai "crawling" is a little bit obscure in this context. so let's rename this class to reflect the fact that this reader only reads the entire content of the sstable. both crawling reader for kl and mx formats are renamed. also, in order to be consistent, all "crawling reader" in variable names are updated as well. --- it's a cleanup, hence no need to backport. Closes scylladb/scylladb#20599 * github.com:scylladb/scylladb: sstable: s/crawling_sstable_mutation_reader/sstable_full_scan_reader sstable/mx/reader: add comment for mx_crawling_sstable_mutation_reader	2024-09-17 11:55:08 +03:00
Kefu Chai	df7f332a58	sstable: s/crawling_sstable_mutation_reader/sstable_full_scan_reader "crawling" is a little bit obscure in this context. so let's rename this class to reflect the fact that this reader only reads the entire content of the sstable. both crawling reader for kl and mx formats are renamed. also, in order to be consistent, all "crawling reader" in variable names are updated as well. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-17 10:39:37 +08:00
Pavel Emelyanov	36863d4ad0	sstables_manager: Remove table_dir from make_sstable() It used to be passed to sstable constructor, but now it doesn't need this argument. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	4425cf54c6	tests: Properly initialize storage options with "dir" Most of the tests work with local storage options. Some support S3 options as well. Whatever it is, when creating an sstable, tests need to put proper "dir" on the options, this patch does so. In fact, storage options for tests are created together with the test-env, and ideally this is the place where dir should be assigned on it. However, there are still places that explicitly specify path they want to see sstables at, for those the new temporary options should be constructed. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:49:50 +03:00
Pavel Emelyanov	56111a50cd	storage_options: Add special-purpose local options maker Lost of code (in tools and tests) explicitly deal with local sstables and need to create options for it. Currently default-constructing options generates local ones, but without the directory path. Add a helper that creates local options with path and patch callers. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-09-13 16:32:39 +03:00
Botond Dénes	c7c5817808	Merge 'Improve timestamp heuristics for tombstone garbage collection' from Benny Halevy When purging regular tombstone consult the min_live_timestamp, if available. This is safe since we don't need to protect dead data from resurrection, as it is already dead. For shadowable_tombstones, consult the min_memtable_live_row_marker_timestamp, if available, otherwise fallback to the min_live_timestamp. If we see in a view table a shadowable tombstone with time T, then in any row where the row marker's timestamp is higher than T the shadowable tombstone is completely ignored and it doesn't hide any data in any column, so the shadowable tombstone can be safely purged without any effect or risk resurrecting any deleted data. In other words, rows which might cause problems for purging a shadowable tombstone with time T are rows with row markers older or equal T. So to know if a whole sstable can cause problems for shadowable tombstone of time T, we need to check if the sstable's oldest row marker (and not oldest column) is older or equal T. And the same check applies similarly to the memtable. If both extended timestamp statistics are missing, fallback to the legacy (and inaccurate) min_timestamp. Fixes scylladb/scylladb#20423 Fixes scylladb/scylladb#20424 > [!NOTE] > no backport needed at this time > We may consider backport later on after given some soak time in master/enterprise > since we do see tombstone accumulation in the field under some materialized views workloads Closes scylladb/scylladb#20446 * github.com:scylladb/scylladb: cql-pytest: add test_compaction_tombstone_gc sstable_compaction_test: add mv_tombstone_purge_test sstable_compaction_test: tombstone_purge_test: test that old deleted data do not inhibit tombstone garbage collection sstable_compaction_test: tombstone_purge_test: add testlog debugging sstable_compaction_test: tombstone_purge_test: make_expiring: use next_timestamp sstable, compaction: add debug logging for extended min timestamp stats compaction: get_max_purgeable_timestamp: use memtable and sstable extended timestamp stats compaction: define max_purgeable_fn tombstone: can_gc_fn: move declaration to compaction_garbage_collector.hh sstables: scylla_metadata: add ext_timestamp_stats compaction_group, storage_group, table_state: add extended timestamp stats getters sstables, memtable: track live timestamps memtable_encoding_stats_collector: update row_marker: do nothing if missing	2024-09-13 08:56:51 +03:00
Takuya ASADA	3cd2a61736	dist: drop scylla-jmx Since JMX server is deprecated, drop them from submodule, build system and package definition. Related scylladb/scylla-tools-java#370 Related #14856 Signed-off-by: Takuya ASADA <syuu@scylladb.com> Closes scylladb/scylladb#17969	2024-09-13 07:59:45 +03:00
Botond Dénes	fc9804ec31	Update tools/java submodule * tools/java 0b4accdd...e505a6d3 (1): > [C-S] Make it use DCAwareRoundRobinPolicy unless rack is provided Closes scylladb/scylladb#20562	2024-09-13 06:30:04 +03:00
Takuya ASADA	90ab2a24df	toolchain: restore multiarch build When we introduced optimized clang at `6e487a4`, we dropped multiarch build on frozen toolchain, because building clang on QEMU emulation is too heavy. Actually, even after the patch merged, there are two mode which does not build clang, --clang-build-mode INSTALL_FROM and --clang-build-mode SKIP. So we should restore multiarch build only these mode, and keep skipping on INSTALL mode since it builds clang. Since we apply multiarch on INSTALL_FROM mode, --clang-archive replaced to --clang-archive-x86_64 and --clang-archive-aarch64. Note that this breaks compatibility of existing clang archive, since it changes clang root directory name from llvm-project to llvm-project-$ARCH. Closes #20442 Closes scylladb/scylladb#20444	2024-09-12 10:44:45 +03:00
Kefu Chai	3e84d43f93	treewide: use seastar::format() or fmt::format() explicitly before this change, we rely on `using namespace seastar` to use `seastar::format()` without qualifying the `format()` with its namespace. this works fine until we changed the parameter type of format string `seastar::format()` from `const char*` to `fmt::format_string<...>`. this change practically invited `seastar::format()` to the club of `std::format()` and `fmt::format()`, where all members accept a templated parameter as its `fmt` parameter. and `seastar::format()` is not the best candidate anymore. despite that argument-dependent lookup (ADT for short) favors the function which is in the same namespace as its parameter, but `using namespace` makes `seastar::format()` more competitive, so both `std::format()` and `seastar::format()` are considered as the condidates. that is what is happening scylladb in quite a few caller sites of `format()`, hence ADT is not able to tell which function the winner in the name lookup: ``` /__w/scylladb/scylladb/mutation/mutation_fragment_stream_validator.cc:265:12: error: call to 'format' is ambiguous 265 \| return format("{} ({}.{} {})", _name_view, s.ks_name(), s.cf_name(), s.id()); \| ^~~~~~ /usr/bin/../lib/gcc/x86_64-redhat-linux/14/../../../../include/c++/14/format:4290:5: note: candidate function [with _Args = <const std::basic_string_view<char> &, const seastar::basic_sstring<char, unsigned int, 15> &, const seastar::basic_sstring<char, unsigned int, 15> &, const utils::tagged_uuid<table_id_tag> &>] 4290 \| format(format_string<_Args...> __fmt, _Args&&... __args) \| ^ /__w/scylladb/scylladb/seastar/include/seastar/core/print.hh:143:1: note: candidate function [with A = <const std::basic_string_view<char> &, const seastar::basic_sstring<char, unsigned int, 15> &, const seastar::basic_sstring<char, unsigned int, 15> &, const utils::tagged_uuid<table_id_tag> &>] 143 \| format(fmt::format_string<A...> fmt, A&&... a) { \| ^ ``` in this change, we change all `format()` to either `fmt::format()` or `seastar::format()` with following rules: - if the caller expects an `sstring` or `std::string_view`, change to `seastar::format()` - if the caller expects an `std::string`, change to `fmt::format()`. because, `sstring::operator std::basic_string` would incur a deep copy. we will need another change to enable scylladb to compile with the latest seastar. namely, to pass the format string as a templated parameter down to helper functions which format their parameters. to miminize the scope of this change, let's include that change when bumping up the seastar submodule. as that change will depend on the seastar change. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-09-11 23:21:40 +03:00
Avi Kivity	ed7d352e7d	Merge 'Validate checksums for uncompressed SSTables' from Nikos Dragazis This PR introduces a new file data source implementation for uncompressed SSTables that will be validating the checksum of each chunk that is being read. Unlike for compressed SSTables, checksum validation for uncompressed SSTables will be active for scrub/validate reads but not for normal user reads to ensure we will not have any performance regression. It consists of: * A new file data source for uncompressed SSTables. * Integration of checksums into SSTable's shareable components. The validation code loads the component on demand and manages its lifecycle with shared pointers. * A new `integrity_check` flag to enable the new file data source for uncompressed SSTables. The flag is currently enabled only through the validation path, i.e., it does not affect normal user reads. * New scrub tests for both compressed and uncompressed SSTables, as well as improvements in the existing ones. * A change in JSON response of `scylla validate-checksums` to report if an uncompressed SSTable cannot be validated due to lack of checksums (no `CRC.db` in `TOC.txt`). Refs #19058. New feature, no backport is needed. Closes scylladb/scylladb#20207 * github.com:scylladb/scylladb: test: Add test to validate SSTables with no checksums tools: Fix typo in help message of scylla validate-checksums sstables: Allow validate_checksums() to report missing checksums test: Add test for concurrent scrub/validate operations test: Add scrub/validate tests for uncompressed SSTables test/lib: Add option to create uncompressed random schemas test: Add test for scrub/validate with file-level corruption test: Check validation errors in scrub tests sstables: Enable checksum validation for uncompressed SSTables sstables: Expose integrity option via crawling mutation readers sstables: Expose integrity option via data_consume_rows() sstables: Add option for integrity check in data streams sstables: Remove unused variable sstables: Add checksum in the SSTable components sstables: Introduce checksummed file data source implementation sstables: Replace assert with on_internal_error	2024-09-11 23:09:45 +03:00
Nikos Dragazis	1f275c71b1	tools: Fix typo in help message of scylla validate-checksums Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:39 +03:00
Nikos Dragazis	5c0a7f706b	sstables: Allow validate_checksums() to report missing checksums Change the return type of `sstable::validate_checksums()` from binary (valid/invalid) to a ternary (valid/invalid/no_checksums). The third status represents uncompressed SSTables without a CRC component (no entry for CRC.db in the TOC). Also, change the JSON response of `sstable validate-checksums` to expose the new status. Replace the boolean value for valid/invalid checksums with an object that contains two boolean keys: one that indicates if the SSTable has checksums, and one that indicates if the checksums are valid or not. The second key is optional and appears only if the SSTable has checksums. Finally, update the documentation to reflect the changes in the API. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2024-09-11 13:12:39 +03:00
Benny Halevy	4de4af954f	sstables: scylla_metadata: add ext_timestamp_stats Store and retrieve the optional extended timestamp statistics (min_live_timestamp and min_live_row_marker_timestamp) in the scylla_metadata component. Note that there is no need for a cluster feature to store those attributes since the scylla_metadata on-disk format is extensible so that old sstables can be read by new versions, seeing the extra stats is missing, and new sstables can be read by old versions that ignore unknown scylla metadata section types. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Benny Halevy	6f202cf48b	compaction_group, storage_group, table_state: add extended timestamp stats getters To return the minimum live timestamp and live row-marker timestamp across a compaction_group, storage_group, or table_state. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-09-10 19:05:57 +03:00
Botond Dénes	fc690a60d8	Update tools/cqlsh submodule * tools/cqlsh 86a280a1...b09bc793 (6): > build(deps): bump actions/download-artifact in /.github/workflows > cqlshlib/test: Add test_formatting.py > cqlshlib/test: Use assertEqual instead of assertEquals > cqlsh.py: Send DESCRIBE statement to server before parsing > cqlsh.py: Fix indentation > cqlsh.py: change shebang to /usr/bin/env python3	2024-09-10 08:11:40 +03:00

1 2 3 4 5 ...

967 Commits