scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-06-01 12:36:56 +00:00

Author	SHA1	Message	Date
Botond Dénes	178c271bf4	readers: make upgrade_to_v2() private The only user is the tests of downgrade_to_v1(), which uses it through mutation source. To avoid any new users popping up, we make it a private method of the latter. In the process the pass-through optimization is dropped, it is not needed for tests anyway.	2022-04-28 14:12:24 +03:00
Botond Dénes	272da51f80	test/lib/mutation_source_test: remove upgrade_to_v2 tests We don't have any upgrade_to_v2() left in production code, so no need to keep testing it. Removing it from this test paves the way for removing it for good (not in this series).	2022-04-28 14:12:24 +03:00
Botond Dénes	7420fb9411	readers: remove v1 forwardable reader No users.	2022-04-28 14:12:24 +03:00
Botond Dénes	f527956cdb	readers: remove v1 empty_reader The only user is row level repair: it is replaced with downgrade_to_v1(make_empty_flat_reader_v2()). The row level reader has lots of downgrade_to_v1() calls, we will deal with these later all at once. Another use is the empty mutation source, this is trivially converted to use the v2 variant.	2022-04-28 14:12:24 +03:00
Botond Dénes	ea37e9c04e	readers: remove v1 delegating_reader The only user is a test, which is hereby converted to use the v2 delegating reader.	2022-04-28 14:12:24 +03:00
Botond Dénes	70d019116f	sstables/kl: make reader impl v2 native The conversion is shallow: the meat of the logic remains v1, fragments are converted to v2 right before being pushed into the buffer. This approach is simple, surgical and is still better then a full upgrade_to_v2().	2022-04-28 14:12:24 +03:00
Botond Dénes	a22b02c801	sstables/kl: return v2 reader from factory methods This just moves the upgrade_to_v2() calls to the other side of said factory methods, preparing the ground for converting the kl reader impl to a native v2 one.	2022-04-28 14:12:24 +03:00
Botond Dénes	4b222e7f37	sstables: move mp_row_consumer_reader_k_l to kl/reader.cc Its only user is in said file, so that is a better place for it.	2022-04-28 14:12:24 +03:00
Botond Dénes	4f77e74bd4	partition_snapshot_reader: convert implementation to native v2 The underlying mutation representation is still v1, so the implementation still has to do conversion. This happens right above the lsa reader component.	2022-04-28 14:12:12 +03:00
Botond Dénes	9c7455825b	mutation_fragment_v2: range_tombstone_change: add minimal_memory_usage()	2022-04-28 14:11:51 +03:00
Botond Dénes	024ceec61e	replica/database: drop_column_family(): drop querier cache entries after waiting for ops Reads (part of operations) running concurrent to `drop_column_family()` can create querier cache entries while we wait for them to finish in `await_pending_ops()`. Move the cache entry eviction to after this, to ensure such entries are also cleaned up before destroying the table object. This moves the `_querier_cache.evict_all_for_table()` from `database::remove()` to `database::drop_column_family()`. With that the former doesn't have to return `future<>` anymore. While at it (changing the signature) also rename `column_family` -> `table`. Also add a regression unit test.	2022-04-28 13:40:13 +03:00
Botond Dénes	4c17da9996	replica/database: finish coroutinizing drop_column_family() Said method was already coroutinized, but only halfway, possibly because of the difficulty in expressing `finally()` with coroutines. We now have `coroutines::as_future()` which makes this easier, so finish the job.	2022-04-28 13:40:13 +03:00
Botond Dénes	9b7550f845	replica/database: make remove(const column_family&) private It has no external users. And it shouldn't have either, tables should be removed via drop_column_family().	2022-04-28 13:40:08 +03:00
Avi Kivity	de0ee13f45	schema_tables: forward-declare user_function and user_aggerates These bring in wasm.hh (though they really shouldn't) and make everyone suffer. Forward declare instead and add missing includes where needed. Closes #10444	2022-04-28 07:22:02 +03:00
Botond Dénes	2c08468fcb	Merge 'Make headers self-contained' from Avi Kivity Minor fixlets to make `ninja dev-headers` pass. Closes #10445 * github.com:scylladb/scylla: readers/from_mutations_v2.hh: make self-contained data_dictionary/storage_options.hh: make self-contained	2022-04-28 07:20:10 +03:00
Avi Kivity	a9812166cd	replica, partition_snapshot_reader, keys: replace boost::any with std::any Reduce #include load by standardizing on std::any. In keys.cc, we just drop the unneeded include. One instance of boost::any remains in config_file, due to a tie-in with other boost components. Closes #10441	2022-04-28 07:18:53 +03:00
Avi Kivity	3a81cb7cc3	readers/from_mutations_v2.hh: make self-contained Due to an inline function, we need the definition of flat_mutation_reader_v2.hh, so include it.	2022-04-27 15:55:16 +03:00
Avi Kivity	28406c2c56	data_dictionary/storage_options.hh: make self-contained Add "seastarx.hh" so sstring works (rather than seastar::sstring).	2022-04-27 15:54:32 +03:00
Avi Kivity	333fdcb3f5	Update tools/java submodule (fix NodeProbe: Malformed IPv6 address at index) * tools/java 9bc83b7a32...a4573759a2 (1): > CASSANDRA-17581 fix NodeProbe: Malformed IPv6 address at index Fixes #10442.	2022-04-27 14:51:47 +03:00
Benny Halevy	e88871f4ec	replica: database: move shard_of implementation to mutation layer We don't need the database to determine the shard of the mutation, only its schema. So move the implementation to the respecive definitions of mutation and frozen_mutation. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes #10430	2022-04-27 14:40:24 +03:00
Nadav Har'El	f6ce7891a5	test/alternator: add test for key length limits DynamoDB limits partition-key length to 2048 bytes and sort-key length to 1024 bytes. Alternator currently has no such limits officially, but if a user tries a key length of over 64 KB, the result will be an "internal server error" as Alternator runs into Scylla's low-level key length limit of 64 KB. In this patch we add (mostly xfailing) tests confirming all the above observations. The tests include extensive comments on what they are testing and why. Some of these tests (specifically, the ones checking what happens above 64 KB) should pass once Alternator is fixed. Other tests - requiring that the limits be exactly what they are in DynamoDB - may either not pass or change in the future, depending on what we decide the limits should be in Alternator. Refs #10347 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes #10438	2022-04-26 18:09:19 +02:00
Raphael S. Carvalho	791403e4bb	sstables: Fix deletion of partial SSTables If SSTable write fails, it will leave a partial sst which contains a temporary TOC in addition to other components partially written. temporary TOC content is written upfront, to allow us from deleting all partial components using the former content if write fails. After commit `e5fc4b6`, partial sst cannot be deleted because deletion procedure is incorrectly assuming all SSTs being deleted unconditionally have TOC, but partial SSTs only have TMP TOC instead. That happens because parent_path() requires all path components to exist due to its usage of fs::path::canonical. The consequence of this is that space of partial files cannot be reclaimed, making it worse for Scylla to recover from ENOSPC, which could happen by selecting a set of files for compaction with higher chance of suceeeding given the free space. This is fixed by only calling parent_path() on TMP TOC, which is guaranteed to exist prior to calling fsync_directory(). Fixes #10410. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2022-04-26 11:00:27 -03:00
Raphael S. Carvalho	0be44b1035	sstables: Fix fsync_directory() fsync_directory() is broken because it's unconditionally performing fsync on parent directory, not on the directory that it was called with. To fix, let's remove wrong parent_path() usage. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2022-04-26 11:00:27 -03:00
Raphael S. Carvalho	ca8f5dcdb7	sstables: Rename dirname() to a more descriptive name dirname() is confusing because if it's called on a directory, parent path is retrieved. By renaming it to parent_path(), it's clearer what the function will do exactly. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2022-04-26 11:00:27 -03:00
Avi Kivity	582802825a	treewide: use system-#include (angle brackets) for seastar Seastar is an external library from Scylla's point of view so we should use the angle bracket #include style. Most of the source follows this, this patch fixes a few stragglers. Also fix cases of #include which reached out to seastar's directory tree directly, via #include "seastar/include/sesatar/..." to just refer to <seastar/...>. Closes #10433	2022-04-26 14:46:42 +03:00
Takuya ASADA	48b6aec16a	scripts: use "out()" function for all capture_output subprocesses On `acaf0bb` we applied out() just for perftune.py because we had issue #10390 with this script. But the issue can happen with other commands too, let's apply it to all commands which uses capture_output. related #10390 Closes #10414	2022-04-26 13:56:52 +03:00
Benny Halevy	01f41630a5	compaction: time_window_compaction_strategy: reset estimated_remaining_tasks when running out of candidates _estimated_remaining_tasks gets updated via get_next_non_expired_sstables -> get_compaction_candidates, but otherwise if we return earlier from get_sstables_for_compaction, it does not get updated and may go out of sync. Refs #10418 (to be closed when the fix reaches branch-4.6) Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes #10419	2022-04-26 11:26:48 +03:00
Benny Halevy	055141fc2e	multishard_mutation_query: do_query: stop ctx if lookup_readers fails lookup_readers might fail after populating some readers and those better be closed before returning the exception. Fixes #10351 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes #10425	2022-04-26 11:11:52 +03:00
Botond Dénes	bf1b6ced3c	Merge "Make storage_service::bootstrap less if-y" from Pavel Emelyanov " The method in question performs node bootstrap in several different modes (regular, replacing, rnbo) and several subsequent if-else branches just duplicate each-other. This set merges them making the code easier to read. " * 'br-less-branchy-bootstrap' of https://github.com/xemul/scylla: storage_service: Remove pointless check in replace-bootstrap storage_service: Generalize wait for range setup storage_service: Merge common if-else branches in bootstrap storage_service: Move tables bootstrap-ON upwards	2022-04-26 10:58:30 +03:00
Raphael S. Carvalho	d79fb9a12f	docs: Update compaction controller doc The doc is being updated to reflect the changes in the commit `d8833de3bb` ("Redefine Compaction Backlog to tame compaction aggressiveness"). Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2022-04-26 10:50:45 +03:00
Gleb Natapov	7f26a8eef5	raft: actively search for a leader if it is not known for a tick duration For a follower to forward requests to a leader the leader must be known. But there may be a situation where a follower does not learn about a leader for a while. This may happen when a node becomes a follower while its log is up-to-date and there are no new entries submitted to raft. In such case the leader will send nothing to the follower and the only way to learn about the current leader is to get a message from it. Until a new entry is added to the raft's log a follower that does not know who the leader is will not be able to add entries. Kind of a deadlock. Note that the problem is specific to our implementation where failure detection is done by an outside module. In vanilla raft a leader sends messages to all followers periodically, so essentially it is never idle. The patch solves this by broadcasting specially crafted append reject to all nodes in the cluster on a tick in case a leader is not known. The leader responds to this message with an empty append request which will cause the node to learn about the leader. For optimisation purposes the patch sends the broadcast only in case there is actually an operation that waits for leader to be known. Fixes #10379	2022-04-25 14:51:22 +02:00
Kamil Braun	5308a7d7a3	raft: server: return immediately from `wait_for_leader` if leader is known `wait_for_leader` may be called when leader is known. There's nothing to wait for in this case.	2022-04-25 12:59:55 +02:00
Benny Halevy	db676e9e4a	replica: database: apply: make sure the schema is synced or throw internal error Currently an exception is thrown in the apply stage when the schema is not synced, but it is too late since returning an error doesn't pinpoint which code path was using an unsync'ed schema so move the check earlier, before _apply_stage is called. We need to make sure the schema is synced earlier when the mutation is applied so call on_internal_error to generate a backtrace in testing and still throw an error in production. Typically storage_proxy::mutate_locally implicitly ensures the schema is synced by making a global_schema_ptr for it. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Message-Id: <20220424110057.3957597-1-bhalevy@scylladb.com>	2022-04-25 12:18:47 +02:00
Pavel Solodovnikov	654e6726d1	service: storage_service: coroutinize `node_ops_abort_thread()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:11:20 +03:00
Pavel Solodovnikov	b27c989e62	service: storage_service: coroutinize `node_ops_abort()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:11:14 +03:00
Pavel Solodovnikov	f7e84c6138	service: storage_service: coroutinize `node_ops_done()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:11:08 +03:00
Pavel Solodovnikov	6936dbea49	service: storage_service: coroutinize `node_ops_update_heartbeat()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:11:04 +03:00
Pavel Solodovnikov	1c03d01927	service: storage_service: coroutinize `force_remove_completion()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:10:58 +03:00
Pavel Solodovnikov	fc1dfb0ae1	service: storage_service: coroutinize `start_leaving()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:10:54 +03:00
Pavel Solodovnikov	0a3a7534d6	service: storage_service: coroutinize `start_sys_dist_ks()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:10:49 +03:00
Pavel Solodovnikov	15ea74e41f	service: storage_service: coroutinize `prepare_to_join()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:10:43 +03:00
Pavel Solodovnikov	c739fad5d6	service: storage_service: coroutinize `removenode_add_ranges()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:10:05 +03:00
Pavel Solodovnikov	e392fdda96	service: storage_service: coroutinize `unbootstrap()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:09:56 +03:00
Pavel Solodovnikov	8fa7f47a74	service: storage_service: coroutinize `get_changed_ranges_for_leaving()` Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2022-04-25 09:09:04 +03:00
Benny Halevy	bcd35af7cf	replica: table: generate_and_propagate_view_updates: pass mutation to make_flat_mutation_reader_from_mutations_v2 With `f5ef687acd` we can consume the single mutation directly, so there's n need to pass it as a vector of size 1. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Message-Id: <20220424103826.3930895-1-bhalevy@scylladb.com>	2022-04-24 22:19:19 +03:00
Avi Kivity	728479a6ea	Merge 'Fix map subscript crashes when map or subscript is null' from Nadav Har'El In the filtering expression "WHERE m[?] = 2", our implementation was buggy when either the map, or the subscript, was NULL (and also when the latter was an UNSET_VALUE). Our code ended up dereferencing null objects, yielding bizarre errors when we were lucky, or crashes when we were less lucky - see examples of both in issues #10361, #10399, #10401. The existing test `test_null.py::test_map_subscript_null` reproduced all these bugs sporadically. In this series we improve the test to reproduce the separate bugs separately, and also reproduce additional problems (like the UNSET_VALUE). We then define both `m[NULL]` and `NULL[2]` to result in NULL instead of the existing undefined (and buggy, and crashing) behavior. This new definition is consistent with our usual SQL-inspired tradition that NULL "wins" in expressions - e.g., `NULL < 2` is also defined as resulting in NULL. However, this decision differs from Cassandra, where `m[NULL]` is considered an error but `NULL[2]` is allowed. We believe that making `m[NULL]` be a NULL instead of an error is more consistent, and moreover - necessary if we ever want to support more complicate expressions like `m[a]`, where the column `a` can be NULL for some rows and non-NULL for others, and it doesn't make sense to return an "invalid query" error in the middle of the scan. Fixes #10361 Fixes #10399 Fixes #10401 Closes #10420 * github.com:scylladb/scylla: expressions: don't dereference invalid map subscript in filter expressions: fix invalid dereference in map subscript evaluation test/cql-pytest: improve tests for map subscripts and nulls	2022-04-24 21:16:10 +03:00
Avi Kivity	a4be927e23	Revert "memtable_list: futurize clear_and_add" This reverts commit `2325c566d9`. It causes a use-after-free of a memtable. Fixes #10421.	2022-04-24 21:09:48 +03:00
Asias He	953af38281	streaming: Allow drop table during streaming Currently, if a table is dropped during streaming, the streaming would fail with no_such_column_family error. Since the table is dropped anyway, it makes more sense to ignore the streaming result of the dropped table, whether it is successful or failed. This allows users to drop tables during node operations, e.g., bootstrap or decommission a node. This is especially useful for the cloud users where it is hard to coordinate between a node operation by admin and user cql change. This patch also fixes a possible user after free issue by not passing the table reference object around. Fixes #10395 Closes #10396	2022-04-24 17:43:20 +03:00
Tzach Livyatan	607ccf0393	Update doc project name to scylla dev Closes #10342	2022-04-24 17:40:54 +03:00
Nadav Har'El	fbb2a41246	expressions: don't dereference invalid map subscript in filter If we have the filter expression "WHERE m[?] = 2", the existing code simply assumed that the subscript is an object of the right type. However, while it should indeed be the right type (we already have code that verifies that), there are two more options: It can also be a NULL, or an UNSET_VALUE. Either of these cases causes the existing code to dereference a non-object as an object, leading to bizarre errors (as in issue #10361) or even crashes (as in issue #10399). Cassandra returns a invalid request error in these cases: "Unsupported unset map key for column m" or "Unsupported null map key for column m". We decided to do things differently: * For NULL, we consider m[NULL] to result in NULL - instead of an error. This behavior is more consistent with other expressions that contain null - for example NULL[2] and NULL<2 both result in NULL as well. Moreover, if in the future we allow more complex expressions, such as m[a] (where a is a column), we can find the subscript to be null for some rows and non-null for other rows - and throwing an "invalid query" in the middle of the filtering doesn't make sense. * For UNSET_VALUE, we do consider this an error like Cassandra, and use the same error message as Cassandra. However, the current implementation checks for this error only when the expression is evaluated - not before. It means that if the scan is empty before the filtering, the error will not be reported and we'll silently return an empty result set. We currently consider this ok, but we can also change this in the future by binding the expression only once (today we do it on every evaluation) and validating it once after this binding. Fixes #10361 Fixes #10399 Signed-off-by: Nadav Har'El <nyh@scylladb.com>	2022-04-24 16:05:34 +03:00

1 2 3 4 5 ...

31056 Commits