scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-04-28 04:06:59 +00:00

Author	SHA1	Message	Date
Botond Dénes	591876b44e	Merge 'sstables: do not reload components of unlinked sstables' from Lakshmi Narayanan Sreethar The SSTable is removed from the reclaimed memory tracking logic only when its object is deleted. However, there is a risk that the Bloom filter reloader may attempt to reload the SSTable after it has been unlinked but before the SSTable object is destroyed. Prevent this by removing the SSTable from the reclaimed list maintained by the manager as soon as it is unlinked. The original logic that updated the memory tracking in `sstables_manager::deactivate()` is left in place as (a) the variables have to be updated only when the SSTable object is actually deleted, as the memory used by the filter is not freed as long as the SSTable is alive, and (b) the `_reclaimed.erase(sst)` is still useful during shutdown, for example, when the SSTable is not unlinked but just destroyed. Fixes https://github.com/scylladb/scylladb/issues/19722 Closes scylladb/scylladb#19717 github.com:scylladb/scylladb: boost/bloom_filter_test: add testcase to verify unlinked sstables are not reloaded sstables: do not reload components of unlinked sstables sstables/sstables_manager: introduce on_unlink method	2024-07-22 12:08:25 +03:00
Laszlo Ersek	680403d2cd	sstables/storage: synch "dst_dir" more leanly in create_links_common() filesystem_storage::create_links_common() runs on directories that generally differ from "_dir", thus, we can't replace its sync_directory() calls with _dir.sync(). We can still use a common (temporary) "opened_directory" object for synching "dst_dir" three times, saving two open and two close operations. This patch is best viewed with "git show -W". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:46:31 +02:00
Laszlo Ersek	0057ee2431	sstables/storage: close previous directory asynchronously upon dir change In "filesystem_storage", change_dir_for_test() and move() replace "_dir" with "opened_directory(new_dir)" using the move assignment operator. Consequently, the file descriptor underlying "_dir" is closed synchronously as a part of object destruction. Expose the async file::close() function through "opened_directory". Introduce filesystem_storage::change_dir() as a common async workhorse for both change_dir_for_test() and move(). In change_dir(), close the old directory asynchronously. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:43:19 +02:00
Laszlo Ersek	6711574646	sstables/storage: futurize change_dir_for_test() Currently change_dir_for_test() is synchronous. Make it return a future, so that we can use async operations in change_dir_for_test() overrides. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:43:19 +02:00
Laszlo Ersek	ef446c4da0	sstables/storage: sync through "opened_directory" in filesystem...::move() Near the end of filesystem_storage::move(), we sync both the old directory, and the new directory, if "delay_commit" is null. At that point, the new directory is just "_dir"; call _dir.sync() instead of sync_directory(). This patch is best viewed with "git show -W". Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:14:46 +02:00
Laszlo Ersek	4d33640481	sstables/storage: sync through "opened_directory" in the "easy" cases Replace sst.sstable_write_io_check(sync_directory, _dir.native()) with _dir.sync(sst._write_error_handler) Also replace the explicit (but still relatively "easy") open_checked_directory() + flush() + flush() operations in filesystem_storage::seal() with two _dir.sync() calls. Because filesystem_storage::create_links_common() is marked "const", we need to declare "_dir" mutable. Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:14:46 +02:00
Laszlo Ersek	2c01171a4d	sstables/storage: introduce "opened_directory" class "filesystem_storage::_dir" is currently of type "std::filesystem::path". Introduce a new class called "opened_directory", and change the type of "_dir" to the new class "opened_directory". "opened_directory" keeps the directory open, and offers synchronization on that open directory (i.e., without having to reopen the directory every time). In subsequent patches, that will be put to use. The opening and closing of the wrapped directory cannot easily be handled explicitly in the "filesystem_storage" member functions. ( Namely, test::store() and test::rewrite_toc_without_scylla_component() -- both in "test/lib/sstable_utils.hh" -- perform "open -> ... -> seal" sequences, and such a sequence may be executed repeatedly. For example, sstable_directory_shared_sstables_reshard_correctly() [test/boost/sstable_directory_test.cc] does just that; it "reopens" the "filesystem_storage" object repeatedly. ) Rather than trying to restrict the order of "filesystem_storage" member function calls, replace the "opened_directory" object with a new one whenever the directory pathname is re-set; namely in filesystem_storage::change_dir_for_test() and filesystem_storage::move(). Signed-off-by: Laszlo Ersek <laszlo.ersek@scylladb.com>	2024-07-19 15:14:46 +02:00
Lakshmi Narayanan Sreethar	31ff69a13c	sstables: do not reload components of unlinked sstables The SSTable is removed from the reclaimed memory tracking logic only when its object is deleted. However, there is a risk that the Bloom filter reloader may attempt to reload the SSTable after it has been unlinked but before the SSTable object is destroyed. Prevent this by removing the SSTable from the reclaimed list maintained by the manager as soon as it is unlinked. The original logic that updated the memory tracking in `sstables_manager::deactivate()` is left in place as (a) the variables have to be updated only when the SSTable object is actually deleted, as the memory used by the filter is not freed as long as the SSTable is alive, and (b) the `_reclaimed.erase(*sst)` is still useful during shutdown, for example, when the SSTable is not unlinked but just destroyed. Fixes #19722 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-19 13:15:57 +05:30
Lakshmi Narayanan Sreethar	dbf22848a8	sstables/sstables_manager: introduce on_unlink method Added a new method, on_unlink() to the sstable_manager. This method is now used by the sstable to notify the manager when it has been unlinked, enabling the manager to update its bookkeeping as required. The on_unlink method doesn't do anything yet but will be updated by the next patch. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-19 13:15:55 +05:30
Avi Kivity	c3b9e64713	Merge 'sstable::open_sstable: pass origin from the writer' from Lakshmi Narayanan Sreethar Pass origin when opening the sstable from the writer and store it in the sstable object. This will make the origin available for the entire write path. Closes scylladb/scylladb#19721 * github.com:scylladb/scylladb: sstables: use _origin in write path sstable::open_sstable: pass and store origin	2024-07-18 19:30:32 +03:00
Avi Kivity	926a02451e	Merge 'sstables/index_reader: abort reading during shutdown' from Lakshmi Narayanan Sreethar This PR adds support for aborting index reads from within `index_consume_entry_context::consume_input` when the server is being stopped. The abort source is now propagated down to the `index_consume_entry_context`, making it available for `consume_input` to check if an abort has been requested. If an abort is detected, `consume_input` will throw an exception to stop the index read operation. Closes scylladb/scylladb#19453 * github.com:scylladb/scylladb: test/boost: test abort behaviour during index read sstables/index_reader: stop consuming index when abort has been requested sstables::index_consume_entry_context: store abort_source sstable: drop old filter only after the new filter is built during rebuild sstables/sstables_manager: store abort_source in sstable_manager replica/database: pass abort_source to database constructor	2024-07-18 19:26:22 +03:00
Calle Wilund	7918ec2e39	sstable: apply error_handler on open exceptions	2024-07-17 09:36:27 +00:00
Lakshmi Narayanan Sreethar	7b58fa2534	sstables: use _origin in write path Now that the origin is available inside the sstable object, no need to pass it to the methods called in the write path. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:44:28 +05:30
Lakshmi Narayanan Sreethar	b762a09dcd	sstable::open_sstable: pass and store origin Pass origin when opening the sstable from the writer and store it in the sstable object. This will make the origin available for the entire write path. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:43:30 +05:30
Lakshmi Narayanan Sreethar	64dadd5ec2	sstables/index_reader: stop consuming index when abort has been requested When an abort is requested, stop further reading of the index file and throw and exception from index_consume_entry_context::process_state(). Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:42:50 +05:30
Lakshmi Narayanan Sreethar	c2524337a2	sstables::index_consume_entry_context: store abort_source Store abort source inside sstables::index_consume_entry_context, so that the next patch can implement cancelling the index read when abort is requested. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:42:50 +05:30
Lakshmi Narayanan Sreethar	587da62686	sstable: drop old filter only after the new filter is built during rebuild sstable::maybe_rebuild_filter_from_index drops the existing filter first and then rebuilds the new filter as the method is only called before the sstable is sealed. But to make the index read abortable, the old filter can be dropped only after the new filter is built so that in case if the index consumer gets aborted, we still have the old filter to write to disk. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:42:47 +05:30
Lakshmi Narayanan Sreethar	6a3e7a5e7a	sstables/sstables_manager: store abort_source in sstable_manager Add a new member that stores the abort_source. This can later be used by the sstables to check if an abort has been requested. Also implement sstables_manager::get_abort_source() that returns a const reference to the abort source. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-07-16 20:36:06 +05:30
Michał Chojnowski	1a8ee69a43	sstables/mx/writer: when creating local_compression, use the sstables's schema, not the writer's There are two schema's associated with a sstable writer: the sstable's schema (i.e. the schema of the table at the time when the sstable object was created), and the writer's schema (equal to the schema of the reader which is feeding into the writer). It's easy to mix up the two and break something as a result. The writer's schema is needed to correctly interpret and serialize the data passing through the writer, and to populate the on-disk metadata about the on-disk schema. The sstables's schema is used to configure some parameters for newly created sstable, such as bloom filter false positive ratio, or compression. The problem fixed by this patch is that the writer was wrongly creating the compressor objects based on its own schema, but using them based based on the sstable's schema the sstable's schema. This patch forces the writer to use the sstable's schema for both.	2024-07-11 12:53:54 +02:00
Michał Chojnowski	d10b38ba5b	sstables/mx/writer: when creating filter, use the sstables's schema, not the writer's There are two schema's associated with a sstable writer: the sstable's schema (i.e. the schema of the table at the time when the sstable object was created), and the writer's schema (equal to the schema of the reader which is feeding into the writer). It's easy to mix up the two and break something as a result. The writer's schema is needed to correctly interpret and serialize the data passing through the writer, and to populate the on-disk metadata about the on-disk schema. The sstables's schema is used to configure some parameters for newly created sstable, such as bloom filter false positive ratio, or compression. The problem fixed by this patch is that the writer was wrongly creating the filter based on its own schema, while the layer outside the writer was interpreting it as if it was created with the sstable's schema. This patch forces the writer to pick the filter's parameters based on the sstable's schema instead.	2024-07-11 12:53:54 +02:00
Michał Chojnowski	a1834efd82	sstables: for i_filter downcasts, use dynamic_cast instead of static_cast As of this patch, those static_casts are actually invalid in some cases (they cast to the wrong type) because of an oversight. A later patch will fix that. But to even write a reliable reproducer for the problem, we must force the invalid casts to manifest as a crash (instead of weird results). This patch both allows writing a reproducer for the bug and serves as a bit of defensive programming for the future.	2024-07-11 12:53:54 +02:00
Kefu Chai	06ba523818	sstable: extract file_writer out `sstables::write()` has multiple overloads, which are defined in `sstables/writer.hh`. two of these overloads are template functions, which have a template parameter named `W`, which has a type constraint requiring it to fulfill the `Writer` concept. but in `types.hh`, when the compiler tries to instantiate the template function with signature of `write(sstable_version_types v, W& out, const T& t)` with `file_writer` as the template parameter of `w`, `file_writer` is only forward-declared using `class file_writer` in the same header file, so this type is still an incomplete type at that moment. that's why the compiler is not able to determine if `file_writer` fulfills the constraint or not. actually, the declaration of `file_writer` is located in `sstables/writer.hh`, which in turn includes `types.hh`. so they form a cyclic dependency. in this change, in order to break this cycle, we extract file_writer out into a separate header file, so that both `sstables/writer.hh` and `sstables/types.hh` can include it. this address the build failure. Fixes scylladb/scylladb#19667 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19669	2024-07-10 23:32:47 +03:00
Patryk Wrobel	a89e3d10af	code-cleanup: add missing header guards The following command had been executed to get the list of headers that did not contain '#pragma once': 'grep -rnw . -e "#pragma once" --include *.hh -L' This change adds missing include guard to headers that did not contain any guard. Signed-off-by: Patryk Wrobel <patryk.wrobel@scylladb.com> Closes scylladb/scylladb#19626	2024-07-09 18:31:35 +03:00
Lakshmi Narayanan Sreethar	c80df8504c	sstables::maybe_rebuild_filter_from_index: log sstable origin Log the sstable origin when its bloom filter is being rebuilt. The origin has to be passed to the method by the caller as it is not available in the sstable object when the filter is rebuilt. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com> Closes scylladb/scylladb#19601	2024-07-04 10:01:23 +03:00
Lakshmi Narayanan Sreethar	21e463b108	sstables/mx/writer: rebuild bloom filters with bad partition estimates The bloom filters are built with partition estimates, as the actual partition count might not be available in all the cases. If the estimate was bad, the bloom filters might end up too large or too small than their optimal sizes. Rebuild such bloom filters with actual partition count before the filter is written to disk and the sstable is sealed. Fixes #19049 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-06-24 12:06:02 +05:30
Lakshmi Narayanan Sreethar	afc90657d6	sstables/mx/writer: add variable to track number of partitions consumed Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-06-24 12:06:02 +05:30
Lakshmi Narayanan Sreethar	fccb1a11e5	sstable: introduce sstable::maybe_rebuild_filter_from_index() Add method sstable::maybe_rebuild_filter_from_index() that rebuilds bloom filters which had bad partition estimates when they were built. The method checks the false positive rate based on the current bitset size against the configured false positive rate to decide whether a filter needs to be rebuilt. If the current false positive rate is within 75% to 125% of the configured false positive rate, the bloom filter will not be rebuilt. Otherwise, the filter will be rebuilt from the index entries. This method should only be called before an SSTable is sealed as the bloom filter is updated in-place. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-06-24 12:06:02 +05:30
Lakshmi Narayanan Sreethar	a7d77f6304	sstable: add method to return filter format for the given sstable version Extract out the filter format computing logic from sstable::read_filter into a separate function. This is done so that the subsequent patches can make use of this function. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-06-24 12:06:01 +05:30
Kefu Chai	c429a8d8ae	sstables: use "me" sstable format by default in `7952200c`, we changed the `selected_format` from `mc` to `me`, but to be backward compatible the cluster starts with "md", so when the nodes in cluster agree on the "ME_SSTABLE_FORMAT" feature, the format selector believes that the node is already using "ME", which is specified by `_selected_format`. even it is actually still using "md", which is specified by `sstable_manager::_format`, as changed by `54d49c04`. as explained above, it was specified to "md" in hope to be backward compatible when upgrading from an existign installation which might be still using "md". but after a second thought, since we are able to read sstables persisted with older formats, this concern is not valid. in other words, `7952200c` introduced a regression which changed the "default" sstable format from `me` to `md`. to address this, we just change `sstable_manager::_format` to "me", so that all sstables are created using "me" format. a test is added accordingly. Fixes #18995 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19293	2024-06-21 12:56:01 +03:00
Avi Kivity	fdc1449392	treewide: rename flat_mutation_reader_v2 to mutation_reader flat_mutation_reader_v2 was introduced in a pair of commits in 2021: `e3309322c3` "Clone flat_mutation_reader related classes into v2 variants" `08b5773c12` "Adapt flat_mutation_reader_v2 to the new version of the API" as a replacement for flat_mutation_reader, using range_tombstone_change instead of range_tombstone to represent represent range tombstones. See those commits for more information. The transition was incremental; the last use of the original flat_mutation_reader was removed in 2022 in commit `026f8cc1e7` "db: Use mutation_partition_v2 in mvcc" In turn, flat_mutation_reader was introduced in 2017 in commit `748205ca75` "Introduce flat_mutation_reader" To transition from a mutation_reader that nested rows within a partition in a separate stream, to a flat reader that streamed partitions and rows in the same stream. Here, we reclaim the original name and rename the awkward flat_mutation_reader_v2 to mutation_reader. Note that mutation_fragment_v2 remains since we still use the original for compatibilty, sometimes. Some notes about the transition: - files were also renamed. In one case (flat_mutation_reader_test.cc), the rename target already existed, so we rename to mutation_reader_another_test.cc. - a namespace 'mutation_reader' with two definitions existed (in mutation_reader_fwd.hh). Its contents was folded into the mutation_reader class. As a result, a few #includes had to be adjusted. Closes scylladb/scylladb#19356	2024-06-21 07:12:06 +03:00
Avi Kivity	185338c8cf	Merge 'Reduce TWCS off-strategy space overhead' from Raphael "Raph" Carvalho Normally, the space overhead for TWCS is 1/N, where is number of windows. But during off-strategy, the overhead is 100% because input sstables cannot be released earlier. Reshaping a TWCS table that takes ~50% of available space can result in system running out of space. That's fixed by restricting every TWCS off-strategy job to 10% of free space in disk. Tables that aren't big will not be penalized with increased write amplification, as all input (disjoint) sstables can still be compacted in a single round. Fixes #16514. Closes scylladb/scylladb#18137 * github.com:scylladb/scylladb: compaction: Reduce twcs off-strategy space overhead to 10% of free space compaction: wire storage free space into reshape procedure sstables: Allow to get free space from underlying storage replica: don't expose compaction_group to reshape task	2024-06-20 18:51:25 +03:00
Kefu Chai	83c6ae10c4	sstables/compress: put type constraints into template type param more compact this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19284	2024-06-14 09:50:55 +03:00
Raphael S. Carvalho	51c7ee889e	sstables: Allow to get free space from underlying storage That will be used in turn to restrict reshape to 10% of available space in underlying storage. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-06-13 12:43:14 -03:00
Botond Dénes	435c01d1e6	sstables: introduce load_metadata() Loads just the metadata components. No validation. Split off from load(), to allow scylla-sstable to partially load an sstable.	2024-06-12 10:46:38 -04:00
Kefu Chai	9318d21a22	sstables: change const_iterator::value_type to uint64_t in general, the value_type of a `const_iterator` is `T` instead of `const T`, what has the const specifier is `reference`. because, when dereferencing an iterator, the value type does not matter any more, as it always a copy. and GCC-14 points this out: ``` /home/kefu/dev/scylladb/sstables/compress.hh:224:13: error: type qualifiers ignored on function return type [-Werror=ignored-qualifiers] 224 \| value_type operator() const { \| ^~~~~~~~~~ /home/kefu/dev/scylladb/sstables/compress.hh:228:13: error: type qualifiers ignored on function return type [-Werror=ignored-qualifiers] 228 \| value_type operator[](ssize_t i) const { \| ^~~~~~~~~~ ``` so, in this change, let's change the value_type to `uint64_t`. please note, it's not typical to return `value_type` from `operator` or `operator[]` of an iterator. but due to the design of segmented_offsets, we cannot return a reference, so let's keep it this way. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#19186	2024-06-09 19:21:16 +03:00
Botond Dénes	37fd568139	sstables/compress.hh: remove unused forward declaration struct compress if forward declared right before its definition. At some point in the past there was probably some code there using it, but now its gone so remove it. Closes scylladb/scylladb#19168	2024-06-09 17:52:05 +03:00
Kefu Chai	e42d83dc46	treewide: include used headers before this change, we rely on `seastar/util/std-compat.hh` to include the used headers provided by stdandard library. this was necessary before we moved to a C++20 compliant standard library implementation. but since Seastar has dropped C++17 support. its `seastar/util/std-compat.hh` is not responsible for providing these headers anymore. so, in this change, we include the used headers directly instead of relying on `seastar/util/std-compat.hh`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#18883	2024-05-27 17:34:38 +03:00
Kefu Chai	5db315930e	sstables: fix a typo in comment: s/Mimicks/Mimics/ this typo was identified by the codespell workflow Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#18781	2024-05-21 12:14:10 +03:00
Botond Dénes	e1c4e6c151	Merge 'sstables_manager: use maintenance scheduling group to run components reload fiber' from Lakshmi Narayanan Sreethar PR https://github.com/scylladb/scylladb/pull/18186 introduced a fiber that reloads reclaimed bloom filters when memory becomes available. Use maintenance scheduling group to run that fiber instead of running it in the main scheduling group. Fixes #18675 Closes scylladb/scylladb#18721 * github.com:scylladb/scylladb: sstables_manager: use maintenance scheduling group to run components reload fiber sstables_manager: add member to store maintenance scheduling group	2024-05-20 16:38:42 +03:00
Lakshmi Narayanan Sreethar	6f58768c46	sstables_manager: use maintenance scheduling group to run components reload fiber Fixes #18675 Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-05-19 15:23:45 +05:30
Lakshmi Narayanan Sreethar	79f6746298	sstables_manager: add member to store maintenance scheduling group Store that maintenance scheduling group inside the sstables_manager. The next patch will use this to run the components reloader fiber. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-05-19 15:23:45 +05:30
Raphael S. Carvalho	715ae689c0	Implement fast streaming for intra-node migration With intra-node migration, all the movement is local, so we can make streaming faster by just cloning the sstable set of leaving replica and loading it into the pending one. This cloning is underlying storage specific, but s3 doesn't support snapshot() yet (th sstables::storage procedure which clone is built upon). It's only supported by file system, with help of hard links. A new generation is picked for new cloned sstable, and it will live in the same directory as the original. A challenge I bumped into was to understand why table refused to load the sstable at pending replica, as it considered them foreign. Later I realized that sharder (for reads) at this stage of migration will point only to leaving replica. It didn't fail with mutation based streaming, because the sstable writer considers the shard -- that the sstable was written into -- as its owner, regardless of what sharder says. That was fixed by mimicking this behavior during loading at pending. test: ./test.py --mode=dev intranode --repeat=100 passes. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-05-16 00:28:47 +02:00
Tomasz Grabiec	c9e6b4dca7	dht: split_range_to_single_shard: Work with static_sharder only In preparation for intra-node tablet migration, to avoid using deprecated sharder APIs. This function is used for generating sstable sharding metadata. For tablets, it is not invoked, so we can safely work with the static sharder. The call site already passes static_sharder only.	2024-05-16 00:28:47 +02:00
Tomasz Grabiec	4d84451cf1	sstables, gdb: Track readers in a linked list For the purpose of scylla-gdb.py command "scylla active-sstables". Before the patch, readers were located by scanning the heap for live objects with vtable pointers corresponding to readers. It was observed that the test scylla_gdb/test_misc.py::test_active_sstables started failing like this: gdb.error: Error occurred in Python: Cannot access memory at address 0x300000000000000 This could be explained by there being a live object on the heap which used to be a reader but now is a different object, and the _sst field contains some other data which is not a pointer. To fix, track readers explicitly in a linked list so that the gdb script can reliably walk readers. Fixes #18618.	2024-05-16 00:28:46 +02:00
Raphael S. Carvalho	0b2ec3063c	sstables: Fix incremental_reader_selector (for range reads) with tablets incremental_reader_selector is the mechanism for incremental comsumption of disjoint sstables on range reads. tablet_sstable_set was implemented, such that selector is efficient with tablets. The problem is selector is vnode addicted and will only consider a given set exhausted when maximum token is reached. With tablets, that means a range read on first tablet of a given shard will also consume other tablets living in the same shard. That results in combined reader having to work with empty sstable readers of tablets that don't intersect with the range of the read. It won't cause extra I/O because the underlying sstables don't intersect with the range of the read. It's only unnecessary CPU work, as it involves creating readers (= allocation), feeding them into combined reader, which will in turn invoke the sstable readers only to realize they don't have any data for that range. With 100k tablets (ranges), and 100 tablets per shard, and ~5 sstables per tablet, there will be this amount of readers (empty or not): (100k * ((100^2 + 100) / 2) * avg_sstable_per_tablet=5) = ~2.5 billions. ~5000 times more readers, it can be quite significant additional cpu work, even though I/O dominates the most in scans. It's an inefficiency that we rather get rid of. The behavior can be observed from logs (there's 1 sstable for each of 4 tablets, but note how readers are created for every single one of them when reading only 1 tablet range): ``` table - make_reader_v2 - range=(-inf, {-4611686018427387905, end}] incremental_reader_selector - create_new_readers(null): selecting on pos {minimum token, w=-1} sstable - make_reader - reader on (-inf, {-4611686018427387905, end}] for sst 3gfx_..._34qn42... that has range [{-9151620220812943033, start},{-4813568684827439727, end}] incremental_reader_selector - create_new_readers(null): selecting on pos {-4611686018427387904, w=-1} sstable - make_reader - reader on (-inf, {-4611686018427387905, end}] for sst 3gfx_..._368nk2... that has range [{-4599560452460784857, start},{-78043747517466964, end}] incremental_reader_selector - create_new_readers(null): selecting on pos {0, w=-1} sstable - make_reader - reader on (-inf, {-4611686018427387905, end}] for sst 3gfx_..._38lj42... that has range [{851021166589397842, start},{3516631334339266977, end}] incremental_reader_selector - create_new_readers(null): selecting on pos {4611686018427387904, w=-1} sstable - make_reader - reader on (-inf, {-4611686018427387905, end}] for sst 3gfx_..._3dba82... that has range [{5065088566032249228, start},{9215673076482556375, end}] ``` Fix is about making sure the tablet set won't select past the supplied range of the read. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Closes scylladb/scylladb#18556	2024-05-14 07:43:22 +03:00
Lakshmi Narayanan Sreethar	4d22c4b68b	sstable_datafile_test: add testcase to test reclaim during reload Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-05-09 19:57:40 +05:30
Lakshmi Narayanan Sreethar	0b061194a7	sstables_manager: reload previously reclaimed components when memory is available When an SSTable is dropped, the associated bloom filter gets discarded from memory, bringing down the total memory consumption of bloom filters. Any bloom filter that was previously reclaimed from memory due to the total usage crossing the threshold, can now be reloaded back into memory if the total usage can still stay below the threshold. Added support to reload such reclaimed filters back into memory when memory becomes available. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-05-09 17:49:22 +05:30
Lakshmi Narayanan Sreethar	f758d7b114	sstables_manager: start a fiber to reload components Start a fiber that gets notified whenever an sstable gets deleted. The fiber doesn't do anything yet but the following patch will add support to reload reclaimed components if there is sufficient memory. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-05-09 17:49:22 +05:30
Lakshmi Narayanan Sreethar	54bb03cff8	sstables: support reloading reclaimed components Added support to reload components from which memory was previously reclaimed as the total memory of reclaimable components crossed a threshold. The implementation is kept simple as only the bloom filters are considered reclaimable for now. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-05-09 17:48:58 +05:30
Lakshmi Narayanan Sreethar	2340ab63c6	sstables_manager: add new intrusive set to track the reclaimed sstables The new set holds the sstables from where the memory has been reclaimed and is sorted in ascending order of the total memory reclaimed. Signed-off-by: Lakshmi Narayanan Sreethar <lakshmi.sreethar@scylladb.com>	2024-05-09 17:48:58 +05:30

1 2 3 4 5 ...

3440 Commits