scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-06-04 14:03:06 +00:00

Author	SHA1	Message	Date
Botond Dénes	2c3c9563b6	scylla-gdb.py: scylla task_histogram: add --filter-tasks option Allowing to include only task objects in the histogram. Leads to histograms with less noise but might exclude potentially important items due to the filtering being inexact.	2022-07-11 16:35:09 +03:00
Botond Dénes	b3fd03f02d	scylla-gdb.py: scylla task_histogram: use histogram class Instead of open-coded alternative. Spares some lines of code and makes future patching easier.	2022-07-11 16:17:29 +03:00
Botond Dénes	46eeb874fc	scylla-gdb.py: scylla-fiber: extract symbol matching logic Into its own class. Soon there will be another user for it.	2022-07-11 16:17:29 +03:00
Botond Dénes	79d5a4cccb	scylla-gdb.py: histogram: add limit feature Limiting the number of printed lines to the top "limit" ones if provided.	2022-07-11 16:08:29 +03:00
Botond Dénes	a80f5f2ba4	scylla-gdb.py: histogram: handle formatting errors Both the content and the formatter method is caller-provided. Mistakes are easy to come by. Instead of aborting the entire operation, just a print an error if an item fails to format.	2022-07-11 16:06:25 +03:00
Botond Dénes	753c1608dd	scylla-gdb.py: intrusive_slist: avoid infinite recursion in __len__() Said method currently uses a list() to iterate over all elements, determining the length. Passing `self` to `list()` will however make call `len()` first, causing infinite recursion.	2022-07-11 15:53:01 +03:00
Botond Dénes	935f9f7ac4	scylla-gdb.py: scylla smp-queues: add --scheduling-group option Allowing for filtering for async work items belonging to a certain scheduling group.	2022-07-11 15:41:29 +03:00
Botond Dénes	fe0f4d4dc1	scylla-gdb.py: scylla smp-queues: add --content switch When present on the command line, the histogram is created over the content of the queues, rather than the number of items in them. It is possible to filter in combination with --content. In particular it can be used to see the content of a single queue when all three of `--to`, `--from` and `--content` is present on the command line.	2022-07-11 15:41:02 +03:00
Botond Dénes	abc77d07d5	scylla-gdb.py: smp-queue: add filtering capability Allow filtering for the from/to cpu (or both). Useful when looking for queues going to a certain CPU.	2022-07-11 15:19:13 +03:00
Botond Dénes	73563d9800	scylla-gdb.py: make scylla smp-queues fast Currently scylla smp-queues has O(count(vobjects)) time complexity as it works by scanning all objects with a vptr and searching them for a pointer to one of the smp message queues. This is very inefficient and unnecessary. It much better to just look at the queues themselves and sum up the number of items in them. This completes in 1-2 seconds on a core where the old algorithm didn't complete in 2h+.	2022-07-11 15:19:13 +03:00
Botond Dénes	966373fcfb	scylla-gdb.py: fix disagreement between std_deque len() and iter() std_deque implementation was broken, with __len__() and __iter__() disagreeing about the size of the container. Turns out both are wrong in certain situations. Fix the iteration logic and re-base both __len__() and __iter__() on the same node iteration code to prevent future disagreements.	2022-07-11 15:19:13 +03:00
Benny Halevy	7e2d2cf1c1	table: snapshot: coroutine::return_exception_ptr Otherwise, we lose the returned exception_ptr type. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes #11000	2022-07-10 17:56:24 +03:00
Nadav Har'El	2437a42b64	Merge 'cql-pytest: add test_permissions.py' from Piotr Sarna This new test suite is expected to gather all kinds of permissions tests - granting, revoking, authorizing, and so on. Right now it contains a single minimal test which ensures that the default superuser can be granted applicable permissions, which they already have anyway. The test suite added in this pull request will also be useful when developing #10633 - permissions for UDF/UDA infrastructure. Closes #10991 * github.com:scylladb/scylla: cql-pytest: add initial permissions test suite cql-pytest: enable CassandraAuthorizer for Scylla and Cassandra	2022-07-10 09:30:10 +03:00
cvybhu	80dda2bb97	cql3: expr: Fix handling reversed types in limits() There was a bug which caused incorrect results of limits() for columns with reversed clustering order. Such columns have reversed_type as their type and this needs to be taken into account when comparing them. It was introduced in `6d943e6cd0`. This commit replaced uses of get_value_comparator with type_of. The difference between them is that get_value_comparator applied ->without_reversed() on the result type. Because the type was reversed, comparisons like 1 < 2 evaluated to false. This caused the test testIndexOnKeyWithReverseClustering to fail, but sadly it wasn't caught by CI because the CI itself has a bug that makes it skip some tests. The test passes now, although it has to be run manually to check that. Fixes: #10918 Signed-off-by: cvybhu <jan.ciolek@scylladb.com> Closes #10994	2022-07-10 09:24:06 +03:00
Nadav Har'El	a7fa29bceb	cross-tree: fix header file self-sufficiency Scylla's coding standard requires that each header is self-sufficient, i.e., it includes whatever other headers it needs - so it can be included without having to include any other header before it. We have a test for this, "ninja dev-headers", but it isn't run very frequently, and it turns out our code deviated from this requirement in a few places. This patch fixes those places, and after it "ninja dev-headers" succeeds again. Fixes #10995 Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes #10997	2022-07-08 12:59:14 +03:00
Avi Kivity	3b20407f25	Merge 'db: Avoid memtable flush latency on schema merge' from Tomasz Grabiec Currently, applying schema mutations involves flushing all schema tables so that on restart commit log replay is performed on top of latest schema (for correctness). The downside is that schema merge is very sensitive to fdatasync latency. Flushing a single memtable involves many syncs, and we flush several of them. It was observed to take as long as 30 seconds on GCE disks under some conditions. This patch changes the schema merge to rely on a separate commit log to replay the mutations on restart. This way it doesn't have to wait for memtables to be flushed. It has to wait for the commitlog to be synced, but this cost is well amortized. We put the mutations into a separate commit log so that schema can be recovered before replaying user mutations. This is necessary because regular writes have a dependency on schema version, and replaying on top of latest schema satisfies all dependencies. Without this, we could get loss of writes if we replay a write which depends on the latest schema on top of old schema. Also, if we have a separate commit log for schema we can delay schema parsing for after the replay and avoid complexity of recognizing schema transactions in the log and invoking the schema merge logic. I reproduced bad behavior locally on my machine with a tired (high latency) SSD disk, load driver remote. Under high load, I saw table alter (server-side part) taking up to 10 seconds before. After the patch, it takes up to 200 ms (50:1 improvement). Without load, it is 300ms vs 50ms. Fixes #8272 Fixes #8309 Fixes #1459 Closes #10333 * github.com:scylladb/scylla: config: Introduce force_schema_commit_log option config: Introduce unsafe_ignore_truncation_record db: Avoid memtable flush latency on schema merge db: Allow splitting initiatlization of system tables db: Flush system.scylla_local on change migration_manager: Do not drop system.IndexInfo on keyspace drop Introduce SCHEMA_COMMITLOG cluster feature frozen_mutation: Introduce freeze/unfreeze helpers for vectors of mutations db/commitlog: Improve error messages in case of unknown column mapping db/commitlog: Fix error format string to print the version db: Introduce multi-table atomic apply()	2022-07-07 16:03:50 +03:00
Benny Halevy	acae3cc223	treewide: stop use of deprecated coroutine::make_exception Convert most use sites from `co_return coroutine::make_exception` to `co_await coroutine::return_exception{,_ptr}` where possible. In cases this is done in a catch clause, convert to `co_return coroutine::exception`, generating an exception_ptr if needed. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Closes #10972	2022-07-07 15:02:16 +03:00
Piotr Sarna	2e61e50e97	cql-pytest: add initial permissions test suite This new test suite is expected to gather all kinds of permissions tests - granting, revoking, authorizing, and so on. Right now it contains a single minimal test which ensures that the default superuser can be granted applicable permissions, which they already have anyway.	2022-07-07 13:45:26 +02:00
Piotr Sarna	1dc116f4dc	cql-pytest: enable CassandraAuthorizer for Scylla and Cassandra In order to be able to test permissions, an authorizer different than AllowAllAuthorizer (default) must be set. CassandraAuthorizer is thus enabled - it works on default user/password pair, so it doesn't introduce any regressions to the test suite.	2022-07-07 13:45:26 +02:00
Avi Kivity	bfc521ee9c	Merge "Activate compaction_throughput_mb_per_sec option" from Pavel E " The option controlls the IO bandwidth of the compaction sched class. It's not set to be 16MB/s, but is unused. This set makes it 0 by default (which means unlimited), live-updateable and plugs it to the seastar sched group IO throttling. branch: https://github.com/xemul/scylla/tree/br-compaction-throttling-3 tests: unit(dev), v2: https://jenkins.scylladb.com/job/releng/job/Scylla-CI/1010/ , v2: manual config update " * 'br-compaction-throttling-3-a' of https://github.com/xemul/scylla: compaction_manager: Add compaction throughput limit updateable_value: Support dummy observing serialized_action: Allow being observer for updateable_value config: Tune the config option	2022-07-07 13:14:07 +03:00
Tomasz Grabiec	6622e3369a	config: Introduce force_schema_commit_log option	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	b8d20335a4	config: Introduce unsafe_ignore_truncation_record The node now refuses to boot if schema tables were truncated. This adds a config option to ignore truncation records as a workaround if user truncated them manually.	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	6b316f267f	db: Avoid memtable flush latency on schema merge Currently, applying schema mutations involves flushing all schema tables so that on restart commit log replay is performed on top of latest schema (for correctness). The downside is that schema merge is very sensitive to fdatasync latency. Flushing a single memtable involves many syncs, and we flush several of them. It was observed to take as long as 30 seconds on GCE disks under some conditions. This patch changes the schema merge to rely on a separate commit log to replay the mutations on restart. This way it doesn't have to wait for memtables to be flushed. It has to wait for the commitlog to be synced, but this cost is well amortized. We put the mutations into a separate commit log so that schema can be recovered before replaying user mutations. This is necessary because regular writes have a dependency on schema version, and replaying on top of latest schema satisfies all dependencies. Without this, we could get loss of writes if we replay a write which depends on the latest schema on top of old schema. Also, if we have a separate commit log for schema we can delay schema parsing for after the replay and avoid complexity of recognizing schema transactions in the log and invoking the schema merge logic. One complication with this change is that replay_position markers are commitlog-domain specific and cannot cross domains. They are recorded in various places which survive node restart: sstables are annotated with the maximum replay position, and they are present inside truncation records. The former annotation is used by "truncate" operation to drop sstables. To prevent old replay positions from being interpreted in the context in the new schema commitlog domain, the change refuses to boot if there are truncation records, and also prohibits truncation of schema tables. The boot sequence needs to know whether the cluster feature associated with this change was enabled on all nodes. Fetaures are stored in system.scylla_local. Because we need to read it before initializing schema tables, the initialization of tables now has to be split into two phases. The first phase initializes all system tables except schema tables, and later we initialize schema tables, after reading stored cluster features. The commitlog domain is switched only when all nodes are upgraded, and only after new node is restarted. This is so that we don't have to add risky code to deal with hot-switching of the commitlog domain. Cold switching is safer. This means that after upgrade there is a need for yet another rolling restart round. Fixes #8272 Fixes #8309 Fixes #1459	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	c5ad05c819	db: Allow splitting initiatlization of system tables We will need some system tables to be initialized earlier in the boot so that system.scylla_local can be read before schema tables are initialized.	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	9b3f96047f	db: Flush system.scylla_local on change So that it can be read before commit log replay. SCHEMA_COMMITLOG feature relies on that.	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	609bf1d547	migration_manager: Do not drop system.IndexInfo on keyspace drop It's not needed anymore because system.IndexInfo is a virtual table calculated from view info. The drop accesses a table which is outside system_schema keyspace so crosses commit log domain. This will trigger an internal from database::apply() on schema merge once the code switches to use the schema commit log and require that all mutations which are part of the schema change belong to a single commit log domain. We could theoretically move system.IndexInfo to the schema commit log domain. It's not easy though because table initialization at boot needs to be split, and current functions for initailization work at keyspace granularity, not table granularity.	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	62df9f446c	Introduce SCHEMA_COMMITLOG cluster feature	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	6c112bf854	frozen_mutation: Introduce freeze/unfreeze helpers for vectors of mutations	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	4eb4689d8c	db/commitlog: Improve error messages in case of unknown column mapping Include the table id, and also add a debug-level log line with replay pos which is similar to the one logged when no error happens.	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	f62eb186b4	db/commitlog: Fix error format string to print the version It always printed {} instead.	2022-07-06 22:08:56 +02:00
Tomasz Grabiec	6444d959dc	db: Introduce multi-table atomic apply() Will be used to apply schema mutations atomically.	2022-07-06 22:08:56 +02:00
Petr Gusev	6cdd5b9ff5	raft, set_configuration fix: don't use dummy entries Leader which ceases to be a leader as a result of a execute_modify_config cannot wait for a dummy record to be committed because io_fiber aborts current waiters as soon as it detects a lost of leadership. This commit excludes dummy entries from the configuration change procedure. A special promise is set on io_fiber when it gets a non-joint configuration, and set_configuration just waits for the corresponding future instead of a dummy record. Fixes: #10010 Closes #10905	2022-07-06 11:26:59 +02:00
Avi Kivity	419fe65259	Revert "Merge 'Block flush until compaction finishes if sstables accumulate' from Mikołaj Sielużycki" This reverts commit `aa8f135f64`, reversing changes made to `9a88bc260c`. The patch causes hangs during flush. Also reverts parts of `411231da75` that impacted the unit test. Fixes #10897.	2022-07-06 12:19:02 +03:00
Asias He	a33c370f9a	gossip: Speed up wait for gossip settle In a large cluster, a node would receive frequent and periodic gossip application state updates like CACHE_HITRATES or VIEW_BACKLOG from peer nodes. Those states are not critical. They should not be counted for the _msg_processing counter which is used to decide if gossip is settled. This patch fixes the long settle on every restart issue reported by users. Refs #10337 Closes #10892	2022-07-06 11:26:32 +03:00
Pavel Emelyanov	b112a98318	compaction_manager: Add compaction throughput limit Re-use eisting compaction_throughput_mb_per_sec option, push it down to compaction manager via config and update the nderlying compaction sched class when the option is (live)updated. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2022-07-06 08:17:08 +03:00
Pavel Emelyanov	b86d11cf67	updateable_value: Support dummy observing An updateable_value() may come without source attached. One of the options how this can happen is if the value sits on a service config. It's a good option to make the config have some default initialization for the option, but in this case observe()ing an option by the service would step on null pointer dereference. Said that, if a value without source is tried to be observed -- assume that it's OK, but the value would never change, so a dummy observer is to be provided. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2022-07-06 08:17:08 +03:00
Pavel Emelyanov	dd0f60ef24	serialized_action: Allow being observer for updateable_value Live-updating an option may involve running some action when the option changes, not just getting its new value into somewhere. The action is nice to be run as serialized action to batch config updates. Said that, here's a sugar to write serialized_action _foo = [this] { return foo(); }; observer<> _o = option.observe(_foo.make_observer()); instead of serialized_action _foo = [this] { return foo(); }; observer<> _o = option.observe([this] { // waited with .join on stop (void)_foo.trigger(); }); Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2022-07-06 08:17:08 +03:00
Mikołaj Grzebieluch	ea3da23f3b	cql3: refactor of native types parsing Native types were parsed directly to data_type, where varchar and text were parsed to utf8_type. To get the name of the type there was a call to the data_type method thus getting the name of the varchar type returns "text". To fix this, added new nonterminal type_unreserved_keyword, which parse native types to their names. It replaced native_or_internal_type in unreserved_function_keyword. unreserved_function_keyword is also used to parse usernames, keyspace names, index names, column identifieres, service levels and role names, so this bug was repaired also in them. Fixes: #10642 Closes #10960	2022-07-05 18:09:17 +02:00
Nadav Har'El	a0ffbf3291	test/cql-pytest: fix test that started failing after error message change Recently a change to Scylla's expression implementation changed the standard error message copied from Cassandra: Cannot execute this query as it might involve data filtering and thus may have unpredictable performance. If you want to execute this query despite the performance unpredictability, use ALLOW FILTERING In the special case where the filter is on the partition key, we changed the message to: Only EQ and IN relation are supported on the partition key (unless you use the token() function or allow filtering) We had a cql-pytest test translated from Cassandra's unit test that checked the old message, and started to fail. Unfortunately nobody noticed because a bug in test.py caused it to stop running these translated unit tests. So in this patch, we trivially fix the test to pass again. Instead of insisting on the old message, we check jsut for the string "allow filtering", in lowercase or uppercase. After this patch, the tests passes as expected on both Scylla and Cassandra. Refs #10918 (this test failing is one of the failures reported there) Refs #10962 (test.py stopped running this test) Signed-off-by: Nadav Har'El <nyh@scylladb.com> Closes #10964	2022-07-05 15:24:20 +03:00
Takuya ASADA	ce87e15ecf	scylla_prepare: fix Exception when SET_NIC_AND_DISKS=no and SET_CLOCKSOURCE=yes We shouldn't call get_tune_mode() when NIC tuning is disabled. fixes #10412 Closes #10959	2022-07-05 14:52:52 +03:00
Takuya ASADA	7501465b7c	scylla_util.py: change debug log directory to /var/tmp/scylla Current debug log is bit difficult to collect in CI, to find the debug log we must know which script caused Exception. Because the filename does not include prefix, and also specified directory is shared with other programs. To make things more easily, let's change debug log directory to /var/tmp/scylla. Closes #10730	2022-07-05 14:49:00 +03:00
Avi Kivity	74b02b9719	Merge 'storage_service: track restore_replica_count' from Benny Halevy This mini-series adds an _async_gate to storage_service that is closed on stop() and it performs restore_replica_count under this gate so it can be orderly waited on in stop() Fixes #10672 Closes #10922 * github.com:scylladb/scylla: storage_service: handle_state_removing: restore_replica_count under _async_gate storage_service: add async_gate for background work	2022-07-05 13:18:59 +03:00
Michał Chojnowski	152eff249c	dbuild: fix --security-opt syntax A recent change added `--security-opt label:disable` to the docker options. There are examples of this syntax on the web, but podman and docker manuals don't mention it and it doesn't work on my machine. Fix it into `--security-opt label=disable`, as described by the manuals. Closes #10965	2022-07-05 10:50:31 +03:00
Avi Kivity	33fe28b0c5	Merge 'commitlog allocation/deletion/flush request rate counters + footprint projection' from Calle Wilund Adds measuring the apparent delta vector of footprint added/removed within the timer time slice, and potentially include this (if influx is greater than data removed) in threshold calculation. The idea is to anticipate crossing usage threshold within a time slice, so request a flush slightly earlier, hoping this will give all involved more time to do their disk work. Obviously, this is very akin to just adjusting the threshold downwards, but the slight difference is that we take actual transaction rate vs. segment free rate into account, not just static footprint. Note: this is a very simplistic version of this anticipation scheme, we just use the "raw" delta for the timer slice. A more sophisiticated approach would perhaps do either a lowpass filtered rate (adjust over longer time), or a regression or whatnot. But again, the default persiod of 10s is something of an eternity, so maybe that is superfluous... Closes #10651 * github.com:scylladb/scylla: commitlog: Add (internal) measurement of byte rates add/release/flush-req commitlog: Add counters for # bytes released/flush requested commitlog: Keep track of last flush high position to avoid double request commitlog: Fix counter descriptor language	2022-07-04 16:26:17 +03:00
Botond Dénes	553538392e	Merge "Improve shutdown logging" from Pavel Emelyanov " On stop there's a rather long log-less gap in the middle of storage_service::drain_on_shutdown(). This set adds log in interesting places and while at it tosses the patched code. refs: #10941 " * 'br-shutdown-logging' of https://github.com/xemul/scylla: batchlog_manager: Add drain and stop logging batchlog_manager: Coroutinize drain and stop batchlog_manager: Drain it with shared future commitlog: Add shutdown message database: Move flushing logging compaction_manager: Add logging around drain compaction_manager: Coroutinize drain storage_service: Sanitize stop_transport()	2022-07-04 13:50:16 +03:00
Pavel Emelyanov	5a4e15f65d	Add .mailmap Google group started replacing sender email with the group email recently. Here's the list of spoiled entries combined from seastar and scylla repos Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Message-Id: <20220701160252.11967-1-xemul@scylladb.com>	2022-07-04 13:44:28 +03:00
Pavel Emelyanov	98ff779676	batchlog_manager: Add drain and stop logging Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2022-07-04 13:42:46 +03:00
Pavel Emelyanov	e2007cd317	batchlog_manager: Coroutinize drain and stop This is not identical change, if drain() resolves with exception we end up skipping the gate closing, but since it's stop why bother Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2022-07-04 13:42:46 +03:00
Pavel Emelyanov	8a03683671	batchlog_manager: Drain it with shared future The .drain() method can be called from several places, each needs to wait for its completion. Now this is achieved with the help of a gate, but there's a simpler way Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2022-07-04 13:42:45 +03:00
Pavel Emelyanov	2e1ec36efd	commitlog: Add shutdown message It happens in database::drain(), we know when it starts after keyspaces are flushed, now it's good to know when it completes Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2022-07-04 13:42:45 +03:00

1 2 3 4 5 ...

31871 Commits