scylladb

Author	SHA1	Message	Date
Piotr Sarna	23057dd186	Merge 'Implement RAFT's leader stepdown extension' from Gleb This series implements leader stepdown extension. See patch 4 for justification for its existence. First three patches either implement cleanups to existing code that future patch will touch or fix bugs that need to be fixed in order for stepdown test to work. * 'raft-leader-stepdown-v3' of github.com:scylladb/scylla-dev: raft: add test for leader stepdown raft: introduce leader stepdown procedure raft: fix replication when leader is not part of current config raft: do not update last election time if current leader is not a part of current configuration raft: move log limiting semaphore into the leader state	2021-03-22 09:45:19 +01:00
Avi Kivity	3c44445c07	Merge "Introduce off-strategy compaction for repair-based bootstrap and replace" from Raphael " Scylla suffers with aggressive compaction after repair-based operation has initiated. That translates into bad latency and slowness for the operation itself. This aggressiveness comes from the fact that: 1) new sstables are immediately added to the compaction backlog, so reducing bandwidth available for the operation. 2) new sstables are in bad shape when integrated into the main sstable set, not conforming to the strategy invariant. To solve this problem, new sstables will be incrementally reshaped, off the compaction strategy, until finally integrated into the main set. The solution takes advantage there's only one sstable per vnode range, meaning sstables generated by repair-based operations are disjoint. NOTE: off-strategy for repair-based decommission and removenode will follow this series and require little work as the infrastructure is introduced in this series. Refs #5226. " * 'offstrategy_v7' of github.com:raphaelsc/scylla: tests: Add unit test for off-strategy sstable compaction table: Wire up off-strategy compaction on repair-based bootstrap and replace table: extend add_sstable_and_update_cache() for off-strategy sstables/compaction_manager: Add function to submit off-strategy work table: Introduce off-strategy compaction on maintenance sstable set table: change build_new_sstable_list() to accept other sstable sets table: change non_staging_sstables() to filter out off-strategy sstables table: Introduce maintenance sstable set table: Wire compound sstable set table: prepare make_reader_excluding_sstables() to work with compound sstable set table: prepare discard_sstables() to work with compound sstable set table: extract add_sstable() common code into a function sstable_set: Introduce compound sstable set reshape: STCS: preserve token contiguity when reshaping disjoint sstables	2021-03-22 10:43:13 +02:00
Gleb Natapov	272cb1c1e6	raft: add test for leader stepdown	2021-03-22 10:31:16 +02:00
Gleb Natapov	9d6bf7f351	raft: introduce leader stepdown procedure Section 3.10 of the PhD describes two cases for which the extension can be helpful: 1. Sometimes the leader must step down. For example, it may need to reboot for maintenance, or it may be removed from the cluster. When it steps down, the cluster will be idle for an election timeout until another server times out and wins an election. This brief unavailability can be avoided by having the leader transfer its leadership to another server before it steps down. 2. In some cases, one or more servers may be more suitable to lead the cluster than others. For example, a server with high load would not make a good leader, or in a WAN deployment, servers in a primary datacenter may be preferred in order to minimize the latency between clients and the leader. Other consensus algorithms may be able to accommodate these preferences during leader election, but Raft needs a server with a sufficiently up-to-date log to become leader, which might not be the most preferred one. Instead, a leader in Raft can periodically check to see whether one of its available followers would be more suitable, and if so, transfer its leadership to that server. (If only human leaders were so graceful.) The patch here implements the extension and employs it automatically when a leader removes itself from a cluster.	2021-03-22 10:28:43 +02:00
Benny Halevy	f562c9c2f3	test: sstable_datafile_test: tombstone_purge_test: use a longer ttl As seen in next-3319 unit testing on jenkins The cell ttl may expire during the test (presuming that the test machine was overloaded), leading to: ``` INFO 2021-03-21 10:05:23,048 [shard 0] compaction - [Compact tests.tombstone_purge 2fcaf680-8a1c-11eb-b1b9-97020c5d261e] Compacting [/jenkins/workspace/scylla-master/next/scylla/testlog/release/scylla-af8644ec-7f07-4ffe-80bf-6703a942e435/la-17-big-Data.db:level=0:origin=, ] INFO 2021-03-21 10:05:23,048 [shard 0] compaction - [Compact tests.tombstone_purge 2fcaf680-8a1c-11eb-b1b9-97020c5d261e] Compacted 1 sstables to []. 4kB to 0 bytes (~0% of original) in 0ms = 0 bytes/s. ~128 total partitions merged to 0. ./test/lib/mutation_assertions.hh(108): fatal error: in "tombstone_purge_test": Mutations differ, expected {table: 'tests.tombstone_purge', key: {'id': alpha, token: -7531858254489963}, mutation_partition: { rows: [ { cont: true, dummy: false, position: { bound_weight: 0, }, 'value': { atomic_cell{1,ts=1616313953,expiry=1616313958,ttl=5} }, }, ] } } ...but got: {table: 'tests.tombstone_purge', key: {'id': alpha, token: -7531858254489963}, mutation_partition: { rows: [ { cont: true, dummy: false, position: { bound_weight: 0, }, 'value': { atomic_cell{DEAD,ts=1616313953,deletion_time=1616313953} }, }, ] } } ``` This corresponds to: ``` 2395 auto mut2 = make_expiring(alpha, ttl); 2396 auto mut3 = make_insert(beta); ... 2399 auto sst2 = make_sstable_containing(sst_gen, {mut2, mut3}); ``` Extend (logical) ttl to 10 seconds to reduce flakiness due to real-time timing. Test: sstable_datafile_test(dev) Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Message-Id: <20210321142931.1226850-1-bhalevy@scylladb.com>	2021-03-21 16:42:00 +02:00
Avi Kivity	1e820687eb	Merge "reader_concurrency_semaphore: limit non-admitted inactive reads" from Botond " Due to bad interaction of recent changes (`913d970` and `4c8ab10`) inctive readers that are not admitted have managed to completely fly under the radar, avoiding any sort of limitation. The reason is that pre-admission the permits don't forward their resource cost to the semaphore, to prevent them possibly blocking their own admission later. However this meant that if such a reader is registered as inactive, it completely avoids the normal resource based eviction mechanism and can accumulate without bounds. The real solution to this is to move the semaphore before the cache and make all reads pass admission before they get started (#4758). Although work has been started towards this, it is still a while until it lands. In the meanwhile this patchset provides a workaround in the form of a new inactive state, which -- like admitted -- causes the permit to forward its cost to the semaphore, making sure these un-admitted inactive reads are accounted for and evicted if there is too much of them. Fixes: #8258 Tests: unit(release), dtest(oppartitions_test.py:TestTopPartitions.test_read_by_gause_key_distribution_for_compound_primary_key_and_large_rows_number) " * 'reader-concurrency-semaphore-limit-inactive-reads/v4' of https://github.com/denesb/scylla: test: mutation_reader_test: add test for permit cleanup test: querier_cache_test: add memory based cache eviction test reader_permit: add inactive state querier: insert(): account immediately evicted querier as resource based eviction reader_concurrency_semaphore: fix clear_inactive_reads() reader_concurrency_semaphore: make inactive_read_handle a weak reference reader_concurrency_semaphore: make evict() noexcept reader_concurrency_semaphore: update out-of-date comments	2021-03-21 16:24:54 +02:00
Nadav Har'El	ab75226626	test/cql-pytest: remove xfail from passing test After commit `0bd201d3ca` ("cql3: Skip indexed column for CK restrictions") fixed issue #7888, the test cassandra_tests/validation/entities/frozen_collections_test.py::testClusteringColumnFiltering began passing, as expected. So we can remove its "xfail" label. Refs #7888. cassandra_tests/validation/entities/frozen_collections_test.py::testClusteringColumnFiltering Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210321080522.1831115-1-nyh@scylladb.com>	2021-03-21 16:02:30 +02:00
Nadav Har'El	10bf2ba60a	cql-pytest: translate Cassandra's reproducers for issue #2962 This is a translation of Cassandra's CQL unit test source file validation/entities/SecondaryIndexOnMapEntriesTest.java into our our cql-pytest framework. This test file checks various features of indexing (with secondary index) individual entries of maps. All these tests pass on Cassandra, but fail on Scylla because of issue #2962 - we do not yet support indexing of the content of unfrozen collections. The failing test currently fail as soon as they try to create the index, with the message: "Cannot create secondary index on non-frozen collection or UDT column v". Refs #2962. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210310124638.1653606-1-nyh@scylladb.com>	2021-03-21 12:30:00 +02:00
Avi Kivity	a78f43b071	Merge 'tracing: fast slow query tracing' from Ivan Prisyazhnyy The set of patches introduces a new tracing mode - `fast slow query tracing`. In this mode, Scylla tracks only tracing sessions and omits all tracing events if the tracing context does not have a `full_tracing` state set. Fixes #2572 Motivation --- We want to run production systems with that option always enabled so we could always catch slow queries without an overhead. The next step is we are gonna optimize further the costs of having tracing enabled to minimize session context handling overhead to allow it to be as transparent for the end-user as possible. Fast tracing mode --- To read the status do $ curl -v http://localhost:10000/storage_service/slow_query To enable fast slow-query tracing $ curl -v --request POST http://localhost:10000/storage_service/slow_query\?fast=true\&enable=true Potential optimizations --- - remove tracing::begin(lazy_eval) - replace tracing::begin(string) for enum to remove copying and memory allocations - merge parameters allocations - group parameters check for trace context - delay formatting - reuse prepared statement shared_ptr instead of both copying it and copying its query Performance --- 100% cache hits --- 1 Core: ``` $ SCYLLA_HOME=/home/sitano.public/Projects/scylla build/release/scylla --smp 1 --cpuset 7 --log-to-syslog 0 --log-to-stdout 1 --default-log-level info --network-stack posix --workdir /home/sitano.public/Projects/scylla --developer-mode 1 --listen-address 0.0.0.0 --api-address 0.0.0.0 --rpc-address 0.0.0.0 --broadcast-rpc-address 172.18.0.1 --broadcast-address 127.0.0.1 ./cassandra-stress write n=100000 no-warmup -pop seq=1..100000 -node 127.0.0.1 -log level=verbose -rate threads=1 -mode native cql3 curl --request POST http://localhost:10000/storage_service/slow_query\?fast\=false\&enable\=false for i in $(seq 5); do taskset -c 2,3,4,5 ./cassandra-stress read duration=5m -pop seq=1..100000 -node 127.0.0.1 -log level=verbose -rate threads=4 throttle=30000/s -mode native cql3 done curl --request POST http://localhost:10000/storage_service/slow_query\?fast\=true\&enable\=true for i in $(seq 5); do taskset -c 2,3,4,5 ./cassandra-stress read duration=5m -pop seq=1..100000 -node 127.0.0.1 -log level=verbose -rate threads=4 throttle=30000/s -mode native cql3 done curl --request POST http://localhost:10000/storage_service/slow_query\?fast\=false\&enable\=true for i in $(seq 5); do taskset -c 2,3,4,5 ./cassandra-stress read duration=5m -pop seq=1..100000 -node 127.0.0.1 -log level=verbose -rate threads=4 throttle=30000/s -mode native cql3 done ``` \| qps \| \| \| -- \| -- \| -- \| -- \| -- \| baseline \| fast, slow \| nofast, slow \| %[1-fastslow/baseline] \| 29,018 \| 26,468 \| 23,591 \| 8.79% \| 28,909 \| 26,274 \| 23,584 \| 9.11% \| 28,900 \| 26,547 \| 23,598 \| 8.14% \| 28,921 \| 26,669 \| 23,596 \| 7.79% \| 28,821 \| 26,385 \| 23,601 \| 8.45% stdev \| 70.24030182 \| 150.9678774 \| 6.670832032 \| avg \| 28,914 \| 26,469 \| 23,594 \| stderr \| 0.24% \| 0.57% \| 0.03% \| %[avg/baseline] \| \| 8.46% \| 18.40% \| 8.46% performance degradation in `fast slow query mode` for pure in-memory workload with minimum traces. 18.40% performance degradation in `original slow query mode` for pure in-memory workload with minimum traces. 0% cache hits --- 1GB memory, 1 Core: $ SCYLLA_HOME=/home/sitano.public/Projects/scylla build/release/scylla --memory 1G --smp 1 --cpuset 7 --log-to-syslog 0 --log-to-stdout 1 --default-log-level info --network-stack posix --workdir /home/sitano.public/Projects/scylla --developer-mode 1 --listen-address 0.0.0.0 --api-address 0.0.0.0 --rpc-address 0.0.0.0 --broadcast-rpc-address 172.18.0.1 --broadcast-address 127.0.0.1 2.4GB, 10000000 keys data: $ ./cassandra-stress write n=10000000 no-warmup -pop seq=1..10000000 -node 127.0.0.1 -log level=verbose -rate threads=4 -mode native cql3 $ curl --request POST http://localhost:10000/storage_service/slow_query\?fast\=true\&enable\=true CASSANDRA_STRESS prepared statements with BYPASS CACHE $ taskset -c 2,3,4,5 ./cassandra-stress read duration=5m -pop seq=1..10000000 -node 127.0.0.1 -log level=verbose -rate threads=4 throttle=30000/s -mode native cql3 20000 reads IOPS, 100MB/s from disk \| qps \| \| \| -- \| -- \| -- \| -- \| -- \| baseline reads \| fast, slow reads \| %[1-fastslow/baseline] \| \| 9,575 \| 9,054 \| 5.44% \| \| 9,614 \| 9,065 \| 5.71% \| \| 9,610 \| 9,066 \| 5.66% \| \| 9,611 \| 9,062 \| 5.71% \| \| 9,614 \| 9,073 \| 5.63% \| stdev \| 16.75410397 \| 6.892024376 \| avg \| 9,605 \| 9,064 \| stderr \| 0.17% \| 0.08% \| %[avg/baseline] \| \| 5.63% \| 5.63% performance degradation in `fast slow query mode` for pure on-disk workload with minimum traces. Closes #8314 * github.com:scylladb/scylla: tracing: fast mode unit test tracing: rest api for lightweight slow query tracing tracing: omit tracing session events and subsessions in fast mode	2021-03-21 12:15:17 +02:00
Dejan Mircevski	318f773d81	types: Unreverse tuple subtype for serialization When a tuple value is serialized, we go through every element type and use it to serialize element values. But an element type can be reversed, which is artificially different from the type of the value being read. This results in a server error due to the type mismatch. Fix it by unreversing the element type prior to comparing it to the value type. Fixes #7902 Tests: unit (dev) Signed-off-by: Dejan Mircevski <dejan@scylladb.com> Closes #8316	2021-03-21 12:07:29 +02:00
Dejan Mircevski	0bd201d3ca	cql3: Skip indexed column for CK restrictions When querying an index table, we assemble clustering-column restrictions for that query by going over the base table token, partition columns, and clustering columns. But if one of those columns is the indexed column, there is a problem; the indexed column is the index table's partition key, not clustering key. We end up with invalid clustering slice, which can cause problems downstream. Fix this by skipping the indexed column when assembling the clustering restrictions. Tests: unit (dev) Fixes #7888 Signed-off-by: Dejan Mircevski <dejan@scylladb.com> Closes #8320	2021-03-21 09:52:06 +02:00
Avi Kivity	58b7f225ab	keys: convert trichotomic comparators to return std::strong_ordering A trichotomic comparator returning an int an easily be mistaken for a less comparator as the return types are convertible. Use the new std::strong_ordering instead. A caller in cql3's update_parameters.hh is also converted, following the path of least resistance. Ref #1449. Test: unit (dev) Closes #8323	2021-03-21 09:30:43 +02:00
Tomasz Grabiec	88a019ba21	Merge "raft: respond with snapshot_reply to send_snapshot RPC" from Kostja Currently send_snapshot is the only two-way RPC used by Raft. However, the sender (the leader) does not look at the receiver's reply, other than checks it's not an error. This has the following issues: - if the follower has a newer term and rejects the snapshot for that reason, the leader will not learn about a newer follower term and will not step down - the send_snapshot message doesn't pass through a single-endpoint fsm::step() and thus may not follow the general Raft rules which apply for all messages. - making a general purpose transport that simply calls fsm::step() for every message becomes impossible. Fix it by actually responding with snapshot_reply to send_snapshot RPC, generating this reply in fsm::step() on the follower, and feeding into fsm::step() on the leader. * scylla-dev/raft-send-snapshot-v2: raft: pass snapshot_reply into fsm::step() raft: respond with snapshot_reply to send_snapshot RPC raft: set follower's next_idx when switching to SNAPSHOT mode raft: set the current leader upon getting InstallSnapshot	2021-03-19 18:13:40 +01:00
Raphael S. Carvalho	64d78eae6a	tests: Add unit test for off-strategy sstable compaction Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2021-03-18 16:56:00 -03:00
Nadav Har'El	0b2cf21932	alternator-test: increase read timeout and avoid retries By default the boto3 library waits up to 60 second for a response, and if got no response, it sends the same request again, multiple times. We already noticed in the past that it retries too many times thus slowing down failures, so in our test configuration lowered the number of retries to 3, but the setting of 60-second-timeout plus 3 retries still causes two problems: 1. When the test machine and the build are extremely slow, and the operation is long (usually, CreateTable or DeleteTable involving multiple views), the 60 second timeout might not be enough. 2. If the timeout is reached, boto3 silently retries the same operation. This retry may fail because the previous one really succeeded at least partially! The symptom is tests which report an error when creating a table which already exists, or deleting a table which dooesn't exist. The solution in this patch is first of all to never do retries - if a query fails on internal server error, or times out, just report this failure immediately. We don't expect to see transient errors during local tests, so this is exactly the right behavior. The second thing we do is to increase the default timeout. If 1 minute was not enough, let's raise it to 5 minutes. 5 minutes should be enough for every operation (famous last words...). Even if 5 minutes is not enough for something, at least we'll now see the timeout errors instead of some wierd errors caused by retrying an operation which was already almost done. Fixes #8135 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210222125630.1325011-1-nyh@scylladb.com>	2021-03-18 18:58:08 +02:00
Botond Dénes	7980140549	test: test_utils: do_check()/do_require(): tone down log to trace They are way too noisy to be at debug level. Signed-off-by: Botond Dénes <bdenes@scylladb.com> Message-Id: <20210318143547.101932-1-bdenes@scylladb.com>	2021-03-18 16:59:59 +02:00
Raphael S. Carvalho	439e9b6fab	table: change build_new_sstable_list() to accept other sstable sets Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2021-03-18 11:47:49 -03:00
Raphael S. Carvalho	6e95860e09	table: change non_staging_sstables() to filter out off-strategy sstables SSTables that are off-strategy should be excluded by this function as it's used to select candidates for regular compaction. So in addition to only returning candidates from the main set, let's also rename it to precisely reflect its behavior. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2021-03-18 11:47:49 -03:00
Raphael S. Carvalho	1e7a444a8b	table: Wire compound sstable set From now own, _sstables becomes the compound set, and _main_sstables refer only to the main sstables of the table. In the near future, maintenance set will be introduced and will also be managed by the compound set. So add_sstable() and on_compaction_completion() are changed to explicitly insert and remove sstables from the main set. By storing compound set in _sstables, functions which used _sstables for creating reader, computing statistics, etc, will not have to be changed when we introduce the maintenance set, so code change is a lot minimized by this approach. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2021-03-18 11:46:06 -03:00
Raphael S. Carvalho	e4b5f5ba33	sstable_set: Introduce compound sstable set This new sstable set implementation is useful for combining operation of multiple sstable sets, which can still be referenced individually via its shared ptr reference. It will be used when maintenance set is introduced in table, so a compound set is required to allow both sets to have their operations efficiently combined. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2021-03-18 11:42:49 -03:00
Botond Dénes	ad02f313dd	test: mutation_reader_test: add test for permit cleanup Check that a permit correctly restores the units on the semaphore in each state it can be destroyed in.	2021-03-18 16:18:22 +02:00
Konstantin Osipov	4afa662d62	raft: respond with snapshot_reply to send_snapshot RPC Raft send_snapshot RPC is actually two-way, the follower responds with snapshot_reply message. This message until now was, however, muted by RPC. Do not mute snapshot_reply any more: - to make it obvious the RPC is two way - to feed the follower response directly into leader's FSM and thus ensure that FSM testing results produced when using a test transport are representative of the real world uses of raft::rpc.	2021-03-18 16:56:42 +03:00
Ivan Prisyazhnyy	f00391af8b	tracing: fast mode unit test Signed-off-by: Ivan Prisyazhnyy <ivan@scylladb.com>	2021-03-18 15:05:09 +02:00
Botond Dénes	c822f0d02a	test: querier_cache_test: add memory based cache eviction test Ensure that the memory consumption of querier cache entries is kept under the limit.	2021-03-18 14:58:21 +02:00
Botond Dénes	594636ebbf	querier: insert(): account immediately evicted querier as resource based eviction `reader_concurrency_semaphore::register_inactive_read()` drops the registered inactive read immediately if there is a resource shortage. This is in effect a resource based eviction, so account it as such in `querier::insert()`.	2021-03-18 14:57:57 +02:00
Botond Dénes	1a337d0ec1	reader_concurrency_semaphore: fix clear_inactive_reads() Broken by the move to an intrusive container (`9cbbf40`), which caused said method to only clear the container but not destroy the inactive reads contained therein. This patch restores the previous behaviour and also adds a call the destructor (to ensure inactive reads are cleaned up under any circumstances), as well as a unit test.	2021-03-18 14:57:57 +02:00
Piotr Sarna	2509b7dbde	Merge 'dht: convert ring_position and decorated_key to std::strong_ordering' from Avi Kivity As #1449 notes, trichotomic comparators returning int are dangerous as they can be mistaken for less comparators. This series converts dht::ring_position and dht::decorated_key, as well as a few closely related downstream types, to return std::strong_ordering. Closes #8225 * github.com:scylladb/scylla: dht: ring_position, decorated_key: convert tri_comparators to std::strong_ordering pager: rephrase misleading comparison check test: total_order_checks: prepare for std::strong_ordering test: mutation_test: prepare merge_container for std::strong_ordering intrusive_array: prepare for std::strong_ordering utils: collection-concepts: prepare for std::strong_ordering	2021-03-18 11:51:54 +01:00
Avi Kivity	a5f17b9a2d	test: total_order_checks: prepare for std::strong_ordering Adjust the total_order_check template to work with comparators returning either int (as a temporary compatibility measure) or std::strong_ordering (for #1449 safety).	2021-03-18 12:40:05 +02:00
Avi Kivity	f0092ae475	test: mutation_test: prepare merge_container for std::strong_ordering The function merge_container() accepts a trichotomic comparator returning an int. As #1449 explains, this is dangerous as it could be mistaken for a less comparator. Switch to std::strong_ordering, but leave a compatible merge_container() in place as it is still needed (even after this series).	2021-03-18 12:40:05 +02:00
Nadav Har'El	4a7d3175e9	test/alternator: make another test faster The slowest test in test_streams.py is test_list_streams_paged. It is meant to test the ListStreams operation with paging. The existing test repeated its test four times, for four different stream types. However, there is no reason to suspect that the ListStreams operation might somehow be different for the four stream types... We already have other tests which create streams of the four types, and uses these streams - we don't need the test for ListStreams to also test creating the four types. By doing this test just once, not four times, we can save around 1.5 seconds of test time. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210318073755.1784349-1-nyh@scylladb.com>	2021-03-18 11:24:18 +01:00
Nadav Har'El	79af728335	test/alternator: make tracing test a bit faster In the test test_tracing.py::test_tracing_all, we do some operations and then need to wait until they appear in the tracing table. The current code used an exponentially-increasing delay during this wait, starting with 0.1 seconds and then doubling the delay until we find what we're looking for. However, it turns out that the delay until the data appears in the table is deliberately chosen by Scylla - and is always around 2 seconds. In this case, an exponential delay is really bad - we will usually wait for around 1 seconds too long after the needed wait of 2 seconds. So in this patch we replace the exponential delay by a constant delay - we wait 0.3 seconds between each retry. This change makes the test test_tracing.py::test_tracing_all finish in a little over 2 seconds, instead of a little over 3 seconds before this patch. We cannot reduce this 2 second time any further unless we make the 2-second tracing delay configurable. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210318000040.1782933-1-nyh@scylladb.com>	2021-03-18 11:24:18 +01:00
Nadav Har'El	4e87f95b42	test/alternator: remove slow and unhelpful test The test test_table.py::test_table_streams_on creates tables with various stream types, and then immediately deletes them without testing anything. This is a slow test (taking almost a full second on my laptop), and is redundant because in test_streams.py we have tests which create tables with streams in the same way - but then actually test that things work with these streams. So this test might as well be removed, and this is what we do in this patch. Removing this test shaves another second from the Alternator test suite's run time. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210317230530.1780849-1-nyh@scylladb.com>	2021-03-18 11:24:18 +01:00
Nadav Har'El	879656e3e0	test/alternator: make a test faster, safer and more correct The test test_condition_expression.py::test_condition_expression_with_forbidden_rmw takes half a second to run (dev build, on my laptop), one of the slowest tests in Alternator's test suite. Part of the reason was that it needlessly set the same table to forbidden_rmw, multiple times. Instead of doing that, we switch to using the test_table_s_forbid_rmw fixture, which is a table like test_table_s but created just once in forbid_rmw mode. The result is a faster test (0.05 seconds instead of 0.5 seconds), but also safer if we ever want to run tests in parallel. It also fixes a bug in the test: At the end of the test, we intended to double-check that although the forbid_rmw table forbids read-modify-write operations, it does allow pure writes. Yet the test did this after clearing the forbid_rmw mode... So after this patch the test verifies this on the forbid_rmw table, as intended. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210317222703.1779992-1-nyh@scylladb.com>	2021-03-18 11:24:18 +01:00
Nadav Har'El	1c2e473e62	test/alternator: make a test faster The test test_condition_expression.py::test_condition_expression_with_permissive_write_isolation Currently takes (on my laptop, dev build) a full two seconds, one of the slowest tests. It is not surprising it is slow - it runs five other tests three times each (for three different write isolation modes), but it doesn't have to be this slow. Before this patch, for each of the five tests we switch the write isolation mode three times, and these switches involve schema changes and are fairly slow. So in this patch we reverse the loop - and switch the write isolation mode to the outer loop. This patch halves the runtime of this test - from two seconds to one. Signed-off-by: Nadav Har'El <nyh@scylladb.com> Message-Id: <20210317221045.1779329-1-nyh@scylladb.com>	2021-03-18 11:24:18 +01:00
Benny Halevy	7862cad669	sstable_set: partitioned_sstable_set: clone: do clone all sstables The existing implementation wrongfully shares _all sstables rather than cloning it. This caused a use-after-free in `repair_meta::do_estimate_partitions_on_local_shard` when traversing a shared sstable_set, during which `table::make_reader_excluding_sstables` erased an entry. The erase should have happened on a cloned copy of the sstable_list, not on a shared copy. The regression was introduced in `c3b8757fa1`. Added a unit test that reproduces the share-on-copy issue for partitioned_stable_set (sstables::sstable_set). Fixes #8274 Test: unit(release, debug) DTest: materialized_views_test.py:TestMaterializedViews.simple_repair_test(debug) Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Reviewed-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Message-Id: <20210317145552.701559-1-bhalevy@scylladb.com>	2021-03-18 11:15:59 +02:00
Nadav Har'El	42169b2eef	Merge 'Alternator: add slow query logging' from Piotr Sarna This series adds slow query logging capability to alternator. Queries which last longer than the specified threshold are logged in `system_traces.node_slow_log` and traced. In order to be better prepared for https://github.com/scylladb/scylla/issues/2572, this series also expands the tracing API to allow custom key-value params and adds a custom `alternator_op` parameter to the slow node log. This information can also be deduced from the tracing session id by consulting the system_traces.events table, but https://github.com/scylladb/scylla/issues/2572 's assumption is that this tracing might not always be available in the future. This series comes with a simple test case which checks if operation logs indeed end up in `system_traces.node_slow_log`. Tests: unit(dev, alternator pytest) manual: verified that no operations are logged if slow query logging is disabled; verified that operations that take less time than the threshold are not logged; verified with test_batch.py::test_batch_write_item_large that a large-enough operation is indeed logged and traced. Fixes #8292 Example trace: ```cql cqlsh> select parameters, duration from system_traces.node_slow_log where start_time=b7a44589-8711-11eb-8053-14c6c5faf955; parameters \| duration ---------------------------------------------------------------------------------------------+---------- {'alternator_op': 'DeleteTable', 'query': '{"TableName": "alternator_Test_1615979572905"}'} \| 75732 ``` Closes #8298 * github.com:scylladb/scylla: alternator: add test for slow query logging alternator: allow enabling slow query logging tracing: allow providing a custom session record param	2021-03-18 11:15:59 +02:00
Piotr Sarna	efe734c575	alternator: add test for slow query logging The test checks whether slow queries are properly logged in the system_traces.node_slow_log system table. The test is deterministic because it uses the threshold of 0ms to qualify a query as slow, which effectively makes all queries "slow enough".	2021-03-17 13:24:26 +01:00
Dejan Mircevski	8db24fc03b	cql3/expr: Handle `IN ?` bound to null Previously, we crashed when the IN marker is bound to null. Throw invalid_request_exception instead. Fixes #8265 Tests: unit (dev) Signed-off-by: Dejan Mircevski <dejan@scylladb.com> Closes #8287	2021-03-17 09:59:22 +02:00
Avi Kivity	972ea9900c	Merge 'commitlog: Make pre-allocation drop O_DSYNC while pre-filling' from Calle Wilund Refs #7794 Iff we need to pre-fill segment file ni O_DSYNC mode, we should drop this for the pre-fill, to avoid issuing flushes until the file is filled. Done by temporarily closing, re-opening in "normal" mode, filling, then re-opening. Closes #8250 * github.com:scylladb/scylla: commitlog: Make pre-allocation drop O_DSYNC while pre-filling commitlog: coroutinize allocate_segment_ex	2021-03-17 09:59:22 +02:00
Tomasz Grabiec	40121621f6	Merge "Kill some get_local_migration_manager() calls" from Pavel Emelyanov There are a bunch of such calls in schema altering statements and there's currently no way to obtain the migration manager for such statements, so a relatively big rework needed. The solution in this set is -- all statements' execute() methods are called with query processor as first argument (now the storage proxy is there), query processor references and provides migration manager for statements. Those statements that need proxy can already get it from the query processor. Afterwards table_helper and thrift code can also stop using the global migration manager instance, since they both have query processor in needed places. While patching them a couple of calls to global storage proxy also go away. The new query processor -> migration manager dependency fits into current start-stop sequence: the migration manager is started early, the query processor is started after it. On stop the query processor remains alive, but the migration manager stops. But since no code currently (should) call get_local_migration_manager() it will _not_ call the query_processor::get_migration_manager() either, so this dangling reference is ugly, but safe. Another option could be to make storage proxy reference migration manager, but this dependency doesn't look correct -- migration manager is higher-level service than the storage proxy is, it is migration manager who currently calls storage proxy, but not the vice versa. * xemul/br-kill-some-migration-managers-2: cql3: Get database directly from query processor thrift: Use query_processor::get_migration_manager() table_helper: Use query_processor::get_migration_manager() cql3: Use query_processor::get_migration_manager() (lambda captures cases) cql3: Use query_processor::get_migration_manager() (alter_type statement) cql3: Use query_processor::get_migration_manager() (trivial cases) query_processor: Keep migration manager onboard cql3: Pass query processor to announce_migration:s cql3: Switch to qp (almost) in schema-altering-stmt cql3: Change execute()'s 1st arg to query_processor	2021-03-17 09:59:22 +02:00
Pavel Solodovnikov	93c565a1bf	raft: allow raft server to start with initial term 0 Prior to the fix there was an assert to check in `raft::server_impl::start` that the initial term is not 0. This restriction is completely artificial and can be lifted without any problems, which will be described below. The only place that is dependent on this corner case is in `server_impl::io_fiber`. Whenever term or vote has changed, they will be both set in `fsm::get_output`. `io_fiber` checks whether it needs to persist term and vote by validating that the term field is set (by actually executing a `term != 0` condition). This particular check is based on an unobvious fact that the term will never be 0 in case `fsm::get_output` saves term and vote values, indicating that they need to be persisted. Vote and term can change independently of each other, so that checking only for term obscures what is happening and why even more. In either case term will never be 0, because: 1. If the term has changed, then it's naturally greater than 0, since it's a monotonically increasing value. 2. If the vote has changed, it means that we received a vote request message. In such case we have already updated our term to the requester's term. Switch to using an explicit optional in `fsm_output` so that a reader don't have to think about the motivation behind this `if` and just checks that `term_and_vote` optional is engaged. Given the motivation described above, the corresponding assert(_fsm->get_current_term() != term_t(0)); in `server_impl::start` is removed. Tests: unit(dev) Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2021-03-17 09:59:21 +02:00
Nadav Har'El	e344f74858	Merge 'logalloc: improve background reclaim shares management' from Avi Kivity The log structured allocator's background reclaimer tries to allocate CPU power proportional to memory demand, but a bug made that not happen. Fix the bug, add some logging, and future-proof the timer. Also, harden the test against overcommitted test machines. Fixes #8234. Test: logalloc_test(dev), 20 concurrent runs on 2 cores (1 hyperthread each) Closes #8281 * github.com:scylladb/scylla: test: logalloc_test: harden background reclain test against cpu overcommit logalloc: background reclaim: use default scheduling group for adjusting shares logalloc: background reclaim: log shares adjustment under trace level logalloc: background reclaim: fix shares not updated by periodic timer	2021-03-17 09:59:21 +02:00
Pavel Emelyanov	1de235f4da	query_processor: Keep migration manager onboard The query processor sits upper than the migration manager, in the services layering, it's started after and (will be) stopped before the migration manager. The migration manager is needed in schema altering statements which are called with query processor argument. They will later get the migration manager from the query processor. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-03-15 19:00:58 +03:00
Avi Kivity	65fea203d2	test: logalloc_test: harden background reclain test against cpu overcommit Use thread CPU time instead of real time to avoid an overcommitted machine from not being able to supply enough CPU for the test.	2021-03-15 13:54:49 +02:00
Alejo Sanchez	88063b6e3e	raft: tests: move common helpers to header Move common test helper functions and data structures to a common helpers.hh header. Signed-off-by: Alejo Sanchez <alejo.sanchez@scylladb.com>	2021-03-15 06:16:58 -04:00
Alejo Sanchez	6139ad6337	raft: tests: move boost tests to tests/raft Move raft boost tests to test/raft directory. Signed-off-by: Alejo Sanchez <alejo.sanchez@scylladb.com>	2021-03-15 06:16:58 -04:00
Calle Wilund	48ca01c3ab	commitlog: Make pre-allocation drop O_DSYNC while pre-filling Refs #7794 Iff we need to pre-fill segment file ni O_DSYNC mode, we should drop this for the pre-fill, to avoid issuing flushes until the file is filled. Done by temporarily closing, re-opening in "normal" mode, filling, then re-opening. v2: * More comment v3: * Add missing flush v4: * comment v5: * Split coroutine and fix into separate patches	2021-03-15 09:35:45 +00:00
Tomasz Grabiec	f2ecb4617e	Merge "raft: implement prevoting stage in leader election" from Gleb This is how PhD explain the need for prevoting stage: One downside of Raft's leader election algorithm is that a server that has been partitioned from the cluster is likely to cause a disruption when it regains connectivity. When a server is partitioned, it will not receive heartbeats. It will soon increment its term to start an election, although it won't be able to collect enough votes to become leader. When the server regains connectivity sometime later, its larger term number will propagate to the rest of the cluster (either through the server's RequestVote requests or through its AppendEntries response). This will force the cluster leader to step down, and a new election will have to take place to select a new leader. Prevoting stage is addressing that. In the Prevote algorithm, a candidate only increments its term if it first learns from a majority of the cluster that they would be willing to grant the candidate their votes (if the candidate's log is sufficiently up-to-date, and the voters have not received heartbeats from a valid leader for at least a baseline election timeout). The Prevote algorithm solves the issue of a partitioned server disrupting the cluster when it rejoins. While a server is partitioned, it won't be able to increment its term, since it can't receive permission from a majority of the cluster. Then, when it rejoins the cluster, it still won't be able to increment its term, since the other servers will have been receiving regular heartbeats from the leader. Once the server receives a heartbeat from the leader itself, it will return to the follower state(in the same term). In our implementation we have "stable leader" extension that prevents spurious RequestVote to dispose an active leader, but AppendEntries with higher term will still do that, so prevoting extension is also required. * scylla-dev/raft-prevote-v5: raft: store leader and candidate state in state variant raft: add boost tests for prevoting raft: implement prevoting stage in leader election raft: reset the leader on entering candidate state raft: use modern unordered_set::contains instead of find in become_candidate	2021-03-12 11:15:51 +01:00
Gleb Natapov	e231186a7b	raft: store leader and candidate state in state variant We already have server state dependant state in fsm, so there is no need to maintain "voters" and "tracker" optionals as well. The upside is that optional and variant sates cannot drift apart now.	2021-03-12 11:12:57 +02:00
Gleb Natapov	e17e7d57bd	raft: add boost tests for prevoting	2021-03-12 11:12:57 +02:00

1 2 3 4 5 ...

1394 Commits