scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-04-25 11:00:35 +00:00

Author	SHA1	Message	Date
Tomasz Grabiec	1a7023c85a	config, tablets: Allow tablets_initial_scale_factor to be a fraction We may want fewer than 1 tablets per shard in large clusters. The per-table option is a fraction, so for consistency, this should be too.	2025-02-19 16:29:08 +01:00
Tomasz Grabiec	7e4a61953d	tablets: load_balancer: Pick initial_scale_factor from config So that it can be live-updated.	2025-02-19 16:29:08 +01:00
Tomasz Grabiec	41789962ef	tablets, load_balancer: Fix and improve logging of resize decisions Resize is no longer only due to avg tablet size. Log avg tablet size as an information, not the reason, and log the true reason for target tablet count.	2025-02-19 16:29:07 +01:00
Tomasz Grabiec	d1ccbee7f9	tablets, load_balancer: Log reason for target tablet count Helps in debugging.	2025-02-19 16:29:07 +01:00
Tomasz Grabiec	029505b179	tablets: load_balancer: Move hints processing to tablet scheduler Hints have common meaning for all strategies, so the logic belongs more to make_sizing_plan(). As a side effect, we can reuse shard capacity computation across tables, which reduces computational complexity from O(tablesnodes) to O(tables DCs + nodes)	2025-02-19 16:29:07 +01:00
Tomasz Grabiec	f1bda8d4c1	tablets: load_balancer: Scale down tablet count to respect per-shard tablet count goal The limit is enforced by controlling average per-shard tablet replica count in a given DC, which is controlled by per-table tablet count. This is effective in respecting the limit on individual shards as long as tablet replicas are distributed evenly between shards. There is no attempt to move tablets around in order to enforce limits on individual shards in case of imbalance between shards. If the average per-shard tablet count exceeds the limit, all tables which contribute to it (have replicas in the DC) are scaled down by the same factor. Due to rounding up to the nearest power of 2, we may overshoot the per-shard goal by at most a factor of 2. If different DCs want different scale factors of a given table, the lowest scale factor is chosen for a given table. The limit is configurable. It's a global per-cluster config which controls how many tablet replicas per shard in total we consider to be still ok. It controls tablet allocator behavior, when choosing initial tablet count. Even though it's a per-node config, we don't support different limits per node. All nodes must have the same value of that config. It's similar in that regard to other scheduler config items like tablets_initial_scale_factor and target_tablet_size_in_bytes.	2025-02-19 16:29:07 +01:00
Tomasz Grabiec	94b5165ac7	tablets: Use scheduler's make_sizing_plan() to decide about tablet count of a new table This makes decisions made by the scheduler consistent with decisions made on table creation, with regard to tablet count. We want to avoid over-allocation of tablets when table is created, which would then be reduced by the scheduler's scaling logic. Not just to avoid wasteful migrations post table creation, but to respect the per-shard goal. To respect the per-shard goal, the algorithm will no longer be as simple as looking at hints, and we want to share the algorithm between the scheduler and initial tablet allocator. So invoke the scheduler to get the tablet count when table is created.	2025-02-19 14:40:07 +01:00
Tomasz Grabiec	dd68c1e526	tablets: load_balancer: Determine desired count from size separately from count from options For debugging purposes. Later we will want to know which rule determined the count.	2025-02-19 14:40:07 +01:00
Tomasz Grabiec	e4c5e2ab55	tablets: load_balancer: Determine resize decision from target tablet count The flow is simpler this way, since the decision cannot now be mismatched with target tablet count.	2025-02-19 14:40:07 +01:00
Tomasz Grabiec	35192e2d6f	tablets: load_balancer: Allow splits even if table stats not available This is in preparation for using the sizing plan during table creation where we never have size stats, and hints are the only determining factor for target tablet count.	2025-02-19 14:40:07 +01:00
Tomasz Grabiec	d3ffea77e6	tablets: load_balancer: Extract make_sizing_plan() Resize plan making will now happen in two stages: 1) Determine desired tablet counts per table (sizing plan) 2) Schedule resize decisions We need intermediate step in the resize plan making, which gives us the planned tablet counts, so that we can plug this part of the algorithm into initial tablet allocation on table construction. We want decisisons made by the scheduler to be consistent with decisions made on table creation. We want to avoid over-allocation of tablets when table is created, which would then be reduced by the scheduler. Not just to avoid wasteful migrations post table creation, but to respect the per-shard goal. To respect the per-shard goal, the algorithm will no longer be as simple as looking at hints, and we want to share the algorithm between the scheduler and initial tablet allocator. Also, this sizing plan will be later plugged into a virtual table for observability.	2025-02-19 14:40:06 +01:00
Tomasz Grabiec	b7e5919fdd	tablets: load_balancer: Simplify resize_urgency_cmp() Logic is preserved since target tablet size is constant for all tables. Dropping d.target_max_tablet_size() will allow us to move it to the load_balancer scope.	2025-02-19 14:39:40 +01:00
Tomasz Grabiec	997007a2df	tablets: load_balancer: Keep config items as instance members It fits preexisting pattern for other config items, and makes the code less cluttered because we don't have to carry config items across calls.	2025-02-19 14:39:39 +01:00
Tomasz Grabiec	8eedb551b5	tablets: network_topology_stragy: Coroutinize calculate_initial_tablets_from_topology() To insert preemption points later.	2025-02-19 14:38:49 +01:00
Tomasz Grabiec	eef18d879c	tablets: load_balancer: Extract get_schema_and_rs() For better readability.	2025-02-19 14:38:49 +01:00
Tomasz Grabiec	9d600dd783	tablets: load_balancer: Drop test_mode tablets_test is now creating proper schema in the database, so test_mode is no longer needed.	2025-02-19 14:38:48 +01:00
Avi Kivity	03ae67f9ea	tablets: load_balancer: don't log decisions to do nothing Demote do-nothing decisions to debug level, but keep them at info if we did decide to do nothing (such as migrate a tablet). Information about more major events (like split/merge) is kept at info level. Once log line that logs node information now also logs the datacenter, which was previously supplied by a log line that is now debug-only. Closes scylladb/scylladb#22783	2025-02-17 11:34:27 +03:00
Benny Halevy	20c6ca2813	tablet_allocator: consider tablet options for resize decision Do not merge tablets if that would drop the tablet_count below the minimum provided by hints. Split tablets if the current tablet_count is less than the minimum tablet count calculated using the table's tablet options. TODO: override min_tablet_count if the tablet count per shard is greater than the maximum allowed. In this case the tables tablet counts should be scaled down proportionally. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2025-02-06 18:43:35 +02:00
Benny Halevy	559f083dc6	tablet_allocator: load_balancer: table_size_desc: keep target_tablet_size as member Rather than target_max_tablet_size. We need both the target as well as max and min tablet sizes, so there is no sense in keeping the max and deriving the target and the minimum for the max value. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2025-02-06 08:59:32 +02:00
Benny Halevy	32c2f7579f	network_topology_strategy: allocate_tablets_for_new_table: consider tablet options Use the keyspace initial_tablets for min_tablet_count, if the latter isn't set, then take the maximum of the option-based tablet counts: - min_tablet_count - and expected_data_size_in_gb / target_tablet_size - min_per_shard_tablet_count (via calculate_initial_tablets_from_topology) If none of the hints produce a positive tablet_count, fall back to calculate_initial_tablets_from_topology * initial_scale. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2025-02-06 08:59:32 +02:00
Botond Dénes	69150f0680	Merge 'Fix edge case issues related to tablet draining ' from Tomasz Grabiec Main problem: If we're draining the last node in a DC, we won't have a chance to evaluate candidates and notice that constraints cannot be satisfied (N < RF). Draining will succeed and node will be removed with replicas still present on that node. This will cause later draining in the same DC to fail when we will have 2 replicas which need relocaiton for a given tablet. The expected behvior is for draining to fail, because we cannot keep the RF in the DC. This is consistent, for example, with what happens when removing a node in a 2-node cluster with RF=2. Fixes #21826 Secondary problem: We allowed tablet_draining transition to be exited with undrained nodes, leaving replicas on nodes in the "left" state. Third problem: We removed DOWN nodes from the candidate node set, even when draining. This is not safe because it may lead to overload. This also makes the "main problem" more likely by extending it to the scenario when the DC is DOWN. The overload part in not a problem in practice currently, since migrations will block on global topology barrier if there are DOWN nodes. Closes scylladb/scylladb#21928 * github.com:scylladb/scylladb: tablets: load_balancer: Fail when draining with no candidate nodes tablets: load_balancer: Ignore skip_list when draining tablets: topology_coordinator: Keep tablet_draining transition if nodes are not drained	2025-01-07 13:04:00 +02:00
Avi Kivity	f3eade2f62	treewide: relicense to ScyllaDB-Source-Available-1.0 Drop the AGPL license in favor of a source-available license. See the blog post [1] for details. [1] https://www.scylladb.com/2024/12/18/why-were-moving-to-a-source-available-license/	2024-12-18 17:45:13 +02:00
Tomasz Grabiec	e732ff7cd8	tablets: load_balancer: Fail when draining with no candidate nodes If we're draining the last node in a DC, we won't have a chance to evaluate candidates and notice that constraints cannot be satisfied (N < RF). Draining will succeed and node will be removed with replicas still present on that node. This will cause later draining in the same DC to fail when we will have 2 replicas which need relocaiton for a given tablet. The expected behvior is for draining to fail, because we cannot keep the RF in the DC. This is consistent, for example, with what happens when removing a node in a 2-node cluster with RF=2. Fixes #21826	2024-12-17 12:14:18 +01:00
Tomasz Grabiec	8718450172	tablets: load_balancer: Ignore skip_list when draining When doing normal load balancing, we can ignore DOWN nodes in the node set and just balance the UP nodes among themselves because it's ok to equalize load just in that set, it improves the situation. It's dangerous to do that when draining because that can lead to overloading of the UP nodes. In the worst case, we can have only one non-drained node in the UP set, which would receive all the tablets of the drained node, doubling its load. It's safer to let the drain fail or stall. This is decided by topology coordinator, currently we will fail (on barrier) and rollback.	2024-12-17 12:14:18 +01:00
Aleksandra Martyniuk	d0cda8ebef	replica: check enabled features in tablet_map_to_mutation Before adding a value to a new column in tablet_map_to_mutation check if the column is supported by the whole cluster. Closes scylladb/scylladb#21941	2024-12-17 07:02:11 +02:00
Botond Dénes	5880a1b90b	Merge 'tasks: add tablet migration virtual task' from Aleksandra Martyniuk In this change, tablet_virtual_task starts supporting tablet migration, in addition to tablet repair. Both tablet operations reuse the same virtual_task because their task data is retrieved similarly. However, it changes nothing from the task manager API users' perspective. They can list running migrations or check their statuses all the same as if migration had its own virtual_task. Users can see running migration tasks - finished tasks are not presented with the task manager API. However, the result of the migration (whether it succeeded or failed) would be presented to users, if they use wait API. If a migration was reverted, it will appear to users as failed. We assume that the migration was reverted, when its destination does not contain a tablet replica. Fixes: https://github.com/scylladb/scylladb/issues/21365. No backport, new feature Closes scylladb/scylladb#21729 * github.com:scylladb/scylladb: test: boost: check migration_task_info in tablet_test.cc replica: add repair related fields to tablet_map_to_mutation test: add tests to check the failed migration virtual tasks test: add tests to check the list of migration virtual tasks test: add tests to check migration virtual tasks status test: topology_tasks: generalize repair task functions service: extend tablet_virtual_task::abort service: extend tablet_virtual_task::wait service: extend tablet_virtual_task::get_status_helper service: extend tablet_virtual_task::contains service: extend tablet_virtual_task::get_stats service: tasks: make get_table_id a method of virtual_task_hint service: tasks: extend virtual_task_hint replica: service: add migration_task_info column to system.tablets locator: extend tablet_task_info to cover migration tasks locator: rename tablet_task_info methods	2024-12-13 10:54:03 +02:00
muthu90tech	e49381119d	locator: topology: use node& instead of node* This change goes thru locator:topology to use node& instead of node* where nullptr is not possible. There are places where the node object is used in unordered_set, in those cases the node is wrapped in std::reference_wrapper. Fixes scylladb/scylladb#20357 Closes scylladb/scylladb#21863	2024-12-12 13:22:55 +01:00
Aleksandra Martyniuk	dee6404aa4	locator: rename tablet_task_info methods	2024-12-11 12:07:36 +01:00
Kefu Chai	8d63d31e57	service: fix a typo in comment s/contraints/constraints/ Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> Closes scylladb/scylladb#21851	2024-12-10 15:58:49 +02:00
Tomasz Grabiec	bf18a17bd6	tablets: scheduler: Fix temporary imbalance in a mixed-capacity cluster on decommission When tablet scheduler drains nodes, it chooses target location based on "badness" metric. Nodes with lowest score are preferred. Before the patch, the score which was used was the number of tablets on that node post-movement. This way we populate least-loaded node first. But this works only if nodes have equal number of shards. If nodes have different capacity, then number of tablets is not a good metric, because we don't aim to equalize per-node count, but per-shard count. We assume that each shard has equal capacity. Because of this bug, during decommission, the nodes with fewer shards would be preferred to receive replicas, which may lead to overloading of those nodes. This imbalance would be later fixed by the normal load balancing logic, but it's still problematic. Fixes #21783 Closes scylladb/scylladb#21860	2024-12-10 14:18:03 +02:00
Raphael S. Carvalho	8344722a26	tests/boost: Add test to verify correctness of balancer decisions during merge Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-04 13:11:11 -03:00
Raphael S. Carvalho	3e518c7b23	service: Co-locate sibling tablets for a table undergoing merge This implements the ability for the balancer to co-locate sibling tablets on the same shard. Co-location is low in priority, so regular load balancer is preferred over it. Previous changes allowed balancer to move co-located sibling tablets together, to not undo the co-location work done so far. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 23:55:43 -03:00
Raphael S. Carvalho	0a6d41305a	service: Make merge of resize plan commutative set_resize_plan() breaks commutativity since it may override the resize plans done earlier, for example, when adding co-location migrations in the DC plan. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 21:47:01 -03:00
Raphael S. Carvalho	48dcefbf45	service: Implement tablet map resize for merge This implements the ability to resize the tablet map for merge if the balancer emits the decision to finalize the merge when all sibling replicas are colocated for a table. But the co-location plan is not implemented in the balancer yet, so this is still not in use. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 21:47:01 -03:00
Raphael S. Carvalho	e00798f1b1	service: Rename topology::transition_state::tablet_split_finalization This transition state will be reused by merge completion, so let's rename it to tablet_resize_finalization. The completion handling path will also be reused, so let's rename functions involved similarly. The old name "tablet split finalization" is deprecated but still recognized and points to the correct transition. Otherwise, the reverse lookup would fail when populating topology system table which last state was split finalization. NOTE: I thought of adding a new tablet_merge_finalization, but it would complicate things since more than one table could be ready for either split or merge, so you need a generic transition state for handling resize completion. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	75a6fe6a75	service: Respect initial_tablet_count if table is in growing mode The initial_tablet_count is respected while the table is in "growing mode". The table implicitly enters this mode when created, since we expect the table to be populated thereafter. We say that a table leaves this mode if it required a split above the initial tablet count. After that, we can rely purely on the average size to say that a table is shrinking and requires merge. This is not perfect and we may want to leave the mode too if we detect the table is shrinking (or even not growing for some significant amount of time), before any split happened. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	5d3b9dba47	service: Wire migration_tablet_set into the load balancer If table is undergoing merge, co-located replicas of sibling tablets will be treated by balancer as if they were a single migration candidate. The reason for that is that the balancer must not undo the co-location work done previously on behalf of merge decision. Sibling tablets will be put in the same migration plan, but note that each tablet is still migrated independently in the state machine. The balancer will exclude both co-located tablets from the candidate list if either haven't finished migration yet. It achieves that by pretending migration of sibling tablets succeeded, allowing it to note that tablets are co-located even though either can still be migrating. The load inversion convergence check also happens after picking a candidate now, since the balancer must be aware that co-located tablets are being migrated together and we want to avoid oscillations. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	fd33e6dfad	service: Introduce sorted_replicas_for_tablet_load() Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	ba633b1da2	service: Introduce alias to per-table candidate map type Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	2791923a21	service: Add replication constraint check variant for migration_tablet_set We have a check that moving a tablet from A to B won't violate replication constraints. The contraints might not be the same for two sibling tablets that have co-located replicas. Example: nodes = {A, B, C, D} tablet1 = {A, B, C} tablet2 = {A, B, D} viable target for {tablet1, B} is D. viable target for {tablet2, B} is C. When co-located replicas share a viable target, then a migration can be emitted to preserve co-location. To allow decommission when co-located replicas don't share a viable target, a skip info will be returned for each tablet, even though that means breaking this co-location. Decommission is higher in priority. Also, doing some preparation for integration of migration_tablet_set. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	308741c9cb	service: Add convergence check variant for migration_tablet_set The load inversion convergence check should be able to know when two tablets are being migrated instead of one, to avoid oscillations. This will be wired when migration_tablet_set is wired. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	ed06b4b1e7	service: Add migration helpers for migration_tablet_set Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	a5db92b9e6	service/tablet_allocator: Introduce migration_tablet_set This new type will allow the load balancer to treat co-located tablets as a single candidate (will treat them as if they were already merged), allowing co-located replicas to be migrated together (in the same migration plan). The type is a variant of global_tablet_id and colocated_tablets (which holds the global_tablet_id of the sibling tablets). It will be eventually wired after some more preparation. It will allow for minimal amount of changes in the balancer code. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	4e20a5eeb1	service: Extract update of node load on migrations Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	e2edcf2c88	service: Extract converge check for intra-node migration This extraction will make it easier later when co-located tablets are introduced in load balancer. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Raphael S. Carvalho	4a0c3ca576	service: Extract erase of tablet replicas from candidate list Intra and inter migration can reuse it. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2024-12-03 20:45:20 -03:00
Kefu Chai	f436edfa22	mutation: remove unused "#include"s these unused includes are identified by clang-include-cleaner. after auditing the source files, all of the reports have been confirmed. please note, because `mutation/mutation.hh` does not include `seastar/coroutine/maybe_yield.hh` anymore, and quite a few source files were relying on this header to bring in the declaration of `maybe_yield()`, we have to include this header in the places where this symbol is used. the same applies to `seastar/core/when_all.hh`. Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>	2024-11-29 14:01:44 +08:00
Asias He	b71a563030	repair: Add core tablet repair scheduler support This adds a new tablet migration kind: repair. It allows tablet repair scheduler to use this migration kind to schedule repair jobs. The current repair scheduler implementation does the following: - A tablet is picked to be repaired when the time since last repair is bigger than a threshold (auto repair mode) or it is requested by user (manual repair mode) - The tablet repair can be scheduled along with tablet migration and rebuild. It runs in the tablet_migration track. - Repair jobs are scheduled in a smart way so that at any point in time, there are no more than configured jobs per shard, which is similar to scylla manager's control. In this patch, both the manual repair and the auto repair are not enabled yet.	2024-11-20 09:42:41 +08:00
Asias He	82a10eca55	tablet_allocator: Introduce stream_weight for tablet_migration_streaming_info The stream_weight for repair migration is set to 2, because it requires more work than just moving the tablet around. The stream_weight for all other migrations are set to 1.	2024-11-19 10:04:41 +08:00
Benny Halevy	4b21cca443	treewide: always allow tablets keyspaces With the tablets feature always enabled (Unless gossip toopology changes are forced), the enable_tablets option now controls only the default for newly created keyspaces. Even when set to `false`, tablets are still enabled as a feature and the user may explicitly enable tablets using `CREATE KEYSPACE <name> WITH tablets = {'enabled': true}` Note: best viewed with `git show -w` Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2024-11-07 13:57:39 +02:00

1 2 3

143 Commits