scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-05-23 08:12:08 +00:00

Author	SHA1	Message	Date
Michał Jadwiszczak	b175f5b97d	db/view/view_building_worker: add more logs when flushing base table Add debug logs around flushing the base table to see how long does it take in case of some stalls in view building. Refs SCYLLADB-1261	2026-05-14 10:23:42 +02:00
Michał Jadwiszczak	e002665aa7	db/system_keyspace: replace `make_view_building_task_mutation()` with mutation builder `system.view_building_tasks` is a single partition table, so it makes more sense to use a mutation builder and generate 1 mutation per group0 command instead of generating multiple mutations.	2026-05-12 21:49:18 +02:00
Piotr Dulikowski	7c2b1ea0b5	Merge 'view_building: fix tombstone_warn_threshold warnings' from Michał Jadwiszczak `system.view_building_tasks` is a single-partition Raft group0 table (pk = `"view_building"`, CK = timeuuid). When `clean_finished_tasks()` deletes hundreds of finished tasks, the physical rows remain in SSTables until compaction. Any subsequent read of the partition counts every column of every tombstoned row as a dead cell, triggering `tombstone_warn_threshold` warnings in large clusters. Two-part fix: 1. Range tombstones instead of row tombstones (commits 2–3) Instead of one row tombstone per finished task, find the minimum alive task UUID (`min_alive_uuid`) and emit a single range tombstone `[before_all, min_alive_uuid)` covering all tasks below that boundary. This reduces the tombstone count significantly and also benefits future compaction. 2. Bounded scan with `min_task_id` (commits 4–6) Even with range tombstones, physical rows remain until compaction and still count as dead cells during reads. The only way to avoid them is to not read them at all. - Add a `min_task_id timeuuid` static column to `system.view_building_tasks`. - On every GC, write `min_task_id = min_alive_uuid` atomically with the range tombstone (same Raft batch). - On reload, read `min_task_id` first using a static-only partition slice (empty `_row_ranges` + `always_return_static_content`): the SSTable reader stops immediately after the static row before processing any clustering tombstones — zero dead cells counted. - Use `AND id >= min_task_id` as a lower bound for the main task scan, skipping all tombstoned rows. The static-only read and the bounded scan are gated on the `VIEW_BUILDING_TASKS_MIN_TASK_ID` cluster feature so mixed-version clusters fall back to the full scan. The issue is not critical, so the fix shouldn't be backported. Fixes SCYLLADB-657 Closes scylladb/scylladb#28929 * github.com:scylladb/scylladb: test/cluster/test_view_building_coordinator: add reproducer for tombstone threshold warning docs: document tombstone avoidance in view_building_tasks view_building: add `task_uuid_generator` to `view_building_task_mutation_builder` view_building: introduce `task_uuid_generator` view_building: store `min_alive_uuid` in view building state view_building: set min_task_id when GC-ing finished tasks view_building: add min_task_id support to view_building_task_mutation_builder view_building: add min_task_id static column and bounded scan to system_keyspace view_building: use range tombstone when GC-ing finished tasks view_building: add range tombstone support to view_building_task_mutation_builder view_building: introduce VIEW_BUILDING_TASKS_MIN_TASK_ID cluster feature	2026-05-12 12:38:25 +03:00
Yaniv Kaul	fdebed5746	db: fix missing format placeholders in log and error messages Fix six format string bugs where arguments were silently dropped: - heat_load_balance.cc: pp value was passed but had no {} placeholder. - commitlog_replayer.cc: column_family_id was passed but table= had no {} placeholder. - view_update_generator.cc: _sstables_with_tables.size() was passed but had no {} placeholder. - view_building_worker.cc: exception pointer was passed but the trailing colon had no {} placeholder. - row_locking.cc: partition key and clustering key were passed in error messages but had no {} placeholders. Signed-off-by: Yaniv Kaul <yaniv.kaul@scylladb.com>	2026-05-10 17:49:50 +03:00
Michał Jadwiszczak	b64f2d2e90	view_building: introduce `task_uuid_generator` With the new `min_alive_uuid` saved in the group0 table, we need to make sure that all new tasks are created with time uuid greater than the value saved in `min_alive_uuid`. This patch introduces the `task_uuid_generator` which ensures that when we are generating multiple tasks in one group0 command, each task will have an unique time uuid and each time uuid will be greater than `min_alive_uuid`.	2026-04-22 09:10:14 +02:00
Avi Kivity	0ae22a09d4	LICENSE: Update to version 1.1 Updated terms of non-commercial use (must be a never-customer).	2026-04-12 19:46:33 +03:00
Piotr Dulikowski	ec0231c36c	Merge 'db/view/view_building_worker: lock staging sstables mutex for all necessary shards when creating tasks' from Michał Jadwiszczak To create `process_staging` view building tasks, we firstly need to collect informations about them on shard0, create necessary mutations, commit them to group0 and move staging sstables objects to their original shards. But there is a possible race after committing the group0 command and before moving the staging sstables to their shards. Between those two events, the coordinator may schedule freshly created tasks and dispatch them to the worker but the worker won't have the sstables objects because they weren't moved yet. This patch fixes the race by holding `_staging_sstables_mutex` locks from all necessary shards when executing `create_staging_sstable_tasks()`. With this, even if the task will be scheduled and dispatched quickly, the worker will wait with executing it until the sstables objects are moved and the locks are released. Fixes SCYLLADB-816 This PR should be backported to all versions containing view building coordinator (2025.4 and newer). Closes scylladb/scylladb#29174 * github.com:scylladb/scylladb: db/view/view_building_worker: fix indentation db/view/view_building_worker: lock staging sstables mutex for necessary shards when creating tasks	2026-04-09 08:37:51 +03:00
Michał Jadwiszczak	9cf94116c2	db/view/view_building_worker: fix indentation	2026-04-07 16:12:04 +02:00
Michał Jadwiszczak	c9aa5bb09c	db/view/view_building_worker: lock staging sstables mutex for necessary shards when creating tasks To create `process_staging` view building tasks, we firstly need to collect informations about them on shard0, create necessary mutations, commit them to group0 and move staging sstables objects to their original shards. But there is a possible race after committing the group0 command and before moving the staging sstables to their shards. Between those two events, the coordinator may schedule freshly created tasks and dispatch them to the worker but the worker won't have the sstables objects because they weren't moved yet. This patch fixes the race by holding `_staging_sstables_mutex` locks from necessary shards when executing `create_staging_sstable_tasks()`. With this, even if the task will be scheduled and dispatched quickly, the worker will wait with executing it until the sstables objects are moved and the locks are released. Fixes SCYLLADB-816	2026-04-07 16:11:45 +02:00
Michał Jadwiszczak	51c164c8d2	view_building_worker: extract starting a new batch to state's method Following the previous commit, a new batch cannot be started if the state was already drained. This commit also adds a check that only one batch is running at a time.	2026-04-07 08:39:05 +02:00
Michał Jadwiszczak	639aa223f3	view_building_worker: distinguish between state's `clear()` and `drain()` While both of this methods do the same (abort current batch, clear data), we can clear the state multiple times during view_building_worker lifetime (for instance when processing base table is changed) but `view_building_worker::state::drain()` should be called only once and after this no other work on the state should be done.	2026-04-07 08:39:05 +02:00
Michał Jadwiszczak	7aea524f52	view_building_worker: lock mutexes before breaking them in `drain()` Not doing this may lead to races like SCYLLADB-844. If some consumer is holding a lock of a mutex and `drain()` is just braking the mutex without locking it beforehand, then the consumer may process its code which should be aborted. An example of the race is SCYLLADB-844, where `work_on_tasks()` is holding `_state._mutex` while it is broken by `drain()`. This causes a new batch is started after the `_state` is cleared.	2026-04-07 08:39:00 +02:00
Michał Jadwiszczak	91c7ac1fb2	view_building_worker: execute drain() once Future changes will require that the view building worker is drained only once per its lifetime.	2026-04-07 08:35:02 +02:00
Botond Dénes	6364e35403	replica/table: add get_tombstone_gc_state() Shorthand for get_compaction_manager().get_shared_tombstone_gc_state().get_tombstone_gc_state().	2026-03-03 14:09:28 +02:00
Pavel Emelyanov	19ea05692c	view_build_worker: Do not switch scheduling groups inside work_on_view_building_tasks The handler appeared back in `c9e710dca3`. In this commit it performed the "core" part of the task -- the do_build_range() method -- inside the streaming sched group. The setup code looks seemingly was copied from the view_builder::do_build_step() method and got the explicit switch of the scheduling group. The switch looks both -- justified and not. On one hand, it makes it explict that the activity runs in the streaming scheduling group. On the other hand, the verb already uses RPC index on 1, which is negotiated to be run in streaming group anyway. On the "third hand", even though being explicit the switch happens too late, as there exists a lot of other activities performed by the handler that seems to also belong to the same scheduling group, but which is not switched into explicitly. By and large, it seems better to avoid the explicit switch and rely on the RPC-level negotiation-based sched group switching. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#28397	2026-02-03 07:00:32 +02:00
Pavel Emelyanov	20a2b944df	view_building_worker: Use its own system_keyspace& reference Some code in the worker need to mess with system_keyspace&. While there's a reference on it from the worker object, it gets one via group0 -> group0_client, which is a bit an overkill. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2026-01-27 14:46:48 +03:00
Michał Jadwiszczak	529cd25c51	db/view/view_building_worker: remove unnnecessary empty lines	2025-12-04 12:52:42 +01:00
Michał Jadwiszczak	4fc5fcaec4	db/view/view_building_worker: fix typo	2025-12-04 12:52:42 +01:00
Michał Jadwiszczak	3253b05ec9	db/view/view_building_worker: avoid creating a copy of tasks map The loop can be converted to use an iterator and avoid creating a copy of the tasks map.	2025-12-04 12:52:41 +01:00
Michał Jadwiszczak	597a2ce5f9	db/view/view_building_worker: wrap conditionally compiled code in a scope The code creates a local variable, so it's better to wrap it in a local scope, to the conditionally compiled variable doesn't pollute the external scope.	2025-12-04 12:52:41 +01:00
Michał Jadwiszczak	a5f19af050	db/view/view_building_worker: remove unnecessary CV broadcast After scylladb/scylladb#26897 was merged, the worker doesn't use the view building state machine CV to manage lifetime of batches, so the broadcast is not needed.	2025-12-04 12:52:41 +01:00
Michał Jadwiszczak	b4fe565f07	db/view/view_building_worker: catch general execption in staging task registrator In case of general exception in `view_building_worker::create_staging_sstable_tasks()`, catch it, print it with error level and sleep 1s before retrying. This will allow for the registrator to retry its work in case of failure and it should be easier to detect any bugs in the method.	2025-12-04 12:52:37 +01:00
Michał Jadwiszczak	fb8cbf1615	db/view/view_building: send coordinator's term in the RPC To avoid case when an old coordinator (which hasn't been stopped yet) dictates what should be done, add raft term to the `work_on_view_building_tasks` RPC. The worker needs to check if the term matches the current term from raft server, and deny the request when the term is bad.	2025-11-25 12:14:05 +01:00
Michał Jadwiszczak	24d69b4005	db/view/view_building_state: replace task's state with `aborted` flag After previous commits, we can drop entire task's state and replace it with single boolean flag, which determines if a task was aborted. Once a task was aborted, it cannot get resurrected to a normal state.	2025-11-25 12:14:04 +01:00
Michał Jadwiszczak	08974e1d50	db/view/view_building_worker: change internal implementation This commit doesn't change the logic behind the view building worker but it changes how the worker is executing view building tasks. Previously, the worker had a state only on shard0 and it was reacting to changes in group0 state. When it noticed some tasks were moved to `STARTED` state, the worker was creating a batch for it on the shard0 state. The RPC call was used only to start the batch and to get its result. Now, the main logic of batch management was moved to the RPC call handler. The worker has a local state on each shard and the state contains: - unique ptr to the batch - set of completed tasks - information for which views the base table was flushed So currently, each batch lives on a shard where it has its work to do exclusively. This eliminates a need to do a synchronization between shard0 and work shard, which was a painful point in previous implementation. The worker still reacts to changes in group0 view building state, but currently it's only used to observe whether any view building tasks was aborted by setting `ABORTED` state. To prepare for further changes to drop the view building task state, the worker ignores `IDLE` and `STARTED` states completely.	2025-11-24 11:12:31 +01:00
Michał Jadwiszczak	6d853c8f11	db/view/view_building_coordinator: change `work_on_tasks` RPC return type During the initial implementation of the view builing coordinator, we decided that if a view building task fails locally on the worker (example reason: view update's target replica is not available), the worker will retry this work instead of reporting a failure to the coordinator. However, we left return type of the RPC, which was telling if a task was finished successfully or aborted. But the worker doesn't need to report that a task was aborted, because it's the coordinator, who decides to abort a task. So, this commit changes the return type to list of UUIDs of completed tasks. Previously length of the returned vector needed to be the same as length of the vector sent in the request. No we can drop this restriction and the RPC handler return list of UUIDs of completed tasks (subset of vector sent in the request). This change is required to drop `STARTED` state in next commits. Since Scylla 2025.4 wasn't released yet and we're going to merge this patch before releasing, no RPC versioning or cluster feature is needed.	2025-11-24 11:12:29 +01:00
Michał Jadwiszczak	4bc6361766	db/view/view_building_worker: support sstables intra-node migration We need to be able to load sstables on the target shard during intra-node tablet migration and to cleanup migrated sstables on the source shard.	2025-11-10 10:36:32 +01:00
Michał Jadwiszczak	c99231c4c2	db/view_building_worker: fix indent	2025-11-10 09:02:16 +01:00
Michał Jadwiszczak	2e8c096930	db/view/view_building_worker: don't organize staging sstables by last token There was a problem with staging sstables after tablet merge. Let's say there were 2 tablets and tablet 1 (lower last token) had an staging sstable. Then a tablet merge occured, so there is only one tablet now (higher last token). But entries in `_staging_sstables`, which are grouped by last token, are never adjusted. Since there shouldn't be thousands of sstables, we can just hold list of sstables per table and filter necessary entries when doing `process_staging` view building task.	2025-11-10 09:02:16 +01:00
Piotr Dulikowski	f76917956c	view_building_worker: access tablet map through erm on sstable discovery Currently, the data returned by `database::get_tables_metadata()` and `database::get_token_metadata()` may not be consistent. Specifically, the tables metadata may contain some tablet-based tables before their tablet maps appear in the token metadata. This is going to be fixed after issue scylladb/scylladb#24414 is closed, but for the time being work around it by accessing the token metadata via `table`->effective_replication_map() - that token metadata is guaranteed to have the tablet map of the `table`. Fixes: scylladb/scylladb#26403 Closes scylladb/scylladb#26588	2025-10-21 00:14:39 +02:00
Botond Dénes	24c6476f73	mutation/mutation_compactor: add tombstone_gc_state to query ctor So tombstones can be purged correctly based on the tombstone gc mode. Currently if repair-mode is used, tombstones are not purged at all, which can lead to purged tombstone being re-replicated to replicas which already purged them via read-repair. This is not a correctness problem, tombstones are not included in data query resutl or digest, these purgable tombstone are only a nuissance for read repair, where they can create extra differences between replicas. Note that for the read repair to trigger, some difference other than in purgable tombstones has to exist, because as mentioned above, these are not included in digets. Fixes: scylladb/scylladb#24332 Closes scylladb/scylladb#26351	2025-10-12 17:48:15 +03:00
Piotr Dulikowski	0b800aab17	Merge 'db/view/view_building_worker: move `discover_existing_staging_sstables()` to the foreground' from Michał Jadwiszczak db/view/view_building_worker: move discover_existing_staging_sstables() to the foreground This patch moves `discover_existing_staging_sstables()` to be executed from main level, instead of running it on the background fiber. This method need to be run only once during the startup to collect existing staging sstables, so there is no need to do it in the background. This change will increase debugability of any further issues related to it (like https://github.com/scylladb/scylladb/issues/26403). Fixes https://github.com/scylladb/scylladb/issues/26417 The patch should be backported to 2025.4 Closes scylladb/scylladb#26446 * github.com:scylladb/scylladb: db/view/view_building_worker: move discover_existing_staging_sstables() to the foreground db/view/view_building_worker: futurize and rename `start_background_fibers()`	2025-10-09 18:24:50 +02:00
Michał Jadwiszczak	8d0d53016c	db/view/view_building_worker: update state again if some batch was finished during the update There was a race between loop in `view_building_worker::run_view_building_state_observer()` and a moment when a batch was finishing its work (`.finally()` callback in `view_building_worker::batch::start()`). State observer waits on `_vb_state_machine.event` CV and when it's awoken, it takes group0 read apply mutex and updates its state. While updating the state, the observer looks at `batch::state` field and reacts to it accordingly. On the other hand, when a batch finishes its work, it sets `state` field to `batch_state::finished` and does a broadcast on `_vb_state_machine.event` CV. So if the batch will execute the callback in `.finally()` while the observer is updating its state, the observer may miss the event on the CV and it will never notice that the batch was finished. This patch fixes this by adding a `some_batch_finished` flag. Even if the worker won't see an event on the CV, it will notice that the flag was set and it will do next iteration. Fixes scylladb/scylladb#26204 Closes scylladb/scylladb#26289	2025-10-09 18:17:22 +02:00
Michał Jadwiszczak	84e4e34d81	db/view/view_building_worker: move discover_existing_staging_sstables() to the foreground This patch moves `discover_existing_staging_sstables()` to be executed from main level, instead of running it on the background fiber. This method need to be run only once during the startup to collect existing staging sstables, so there is no need to do it in the background. This change will increase debugability of any further issues related to it (like scylladb/scylladb#26403). Fixes scylladb/scylladb#26417	2025-10-08 11:16:07 +02:00
Michał Jadwiszczak	575dce765e	db/view/view_building_worker: futurize and rename `start_background_fibers()` Next commit will move `discover_existing_staging_sstables()` to the foreground, so to prepare for this we need to futurize `start_background_fibers()` method and change its name to better reflect its purpose.	2025-10-08 10:19:41 +02:00
Piotr Dulikowski	5e5a3c7ec5	view_building_worker.cc: fix spelling (commiting -> committing) The typo is reported by GitHub action on each PR, so let's fix it to reduce the noise for everybody. Closes scylladb/scylladb#26329	2025-09-30 16:47:03 +03:00
Pavel Emelyanov	f6860d1de0	Merge 'mv: run view building worker fibers in streaming group' from Piotr Dulikowski The background fibers of the view building worker are indirectly spawned by the main function, thus the fibers inherit the "main" scheduling group. The main scheduling group is not supposed to be used for regular work, only for initialization and deinitialization, so this is wrong. Wrap the call to `start_backgroud_fibers()` with `with_scheduling_group` and use the streaming scheduling group. The view building worker already handles RPCs in the streaming scheduling group (which do most of the work; background fibers only do some maintenance), so this seems like a good fit. No need to backport, view build coordinator is not a part of any release yet. Closes scylladb/scylladb#26122 * github.com:scylladb/scylladb: mv: fix typo in start_backgroud_fibers mv: run view building worker fibers in streaming group	2025-09-22 15:28:38 +03:00
Piotr Dulikowski	fb0e5784e4	mv: fix typo in start_backgroud_fibers Letter "n" was missing in this name.	2025-09-18 15:50:16 +02:00
Piotr Dulikowski	4ed045a15c	Merge 'db/view/view_building_worker: wrap `shared_sstable` in `foreign_ptr`' from Michał Jadwiszczak When a staging sstable is registered to view building worker, it needs to make a round trip from its original shard to shard 0 (in order to create a view building task) and back (to be eventually processed). Until now this was done using plain `sstables::shared_sstable` (= `lw_shared_ptr`) which is not safe to be moved between shards. This patch fixes this by wrapping the pointer in `foreign_ptr` and obtains necessary informations (owner shard, last token) on the original shard (instead of on shard0). Then all of those objects are put into freshly introduced structure `staging_sstable_task_info`, which can be safely moved between shards. Fixes https://github.com/scylladb/scylladb/issues/25859 View building coordinator isn't present in any release yet, no backport needed. Closes scylladb/scylladb#25832 * github.com:scylladb/scylladb: db/view/view_building_worker: fix indent db/view/view_building_worker: wrap `shared_sstable` in `foreign_ptr` db/view/view_building_worker: use table id in `register_staging_sstable_tasks()` db/view/view_building_worker: move helper functions higher	2025-09-18 10:24:27 +02:00
Michał Jadwiszczak	1d8b41a51d	db/view/view_building_worker: fix indents	2025-09-18 03:24:43 +02:00
Michał Jadwiszczak	99db5a6c30	db/view/view_building_worker: change `sharded<abort_source>` to local `abort_source` Previously the sharded abort_sources was stopped at the end of batch::do_work(), which is working in parallel to view building worker main loop. This leads to races because the worker may call batch::abort(), which access the abort_sources. This patch solves this be changing `sharded<abort_source>` into `abort_source`. Since now `batch::do_work()` is executed on tasks' shard, all abort source checks are also done on tasks' shard. The only place where shard0 uses the abort source is `batch::abort()`, but this method now does `smp::submit_to(replica.shard, [request abort])`, so the abort source is used on tasks' shard exclusively. Fixes scylladb/scylladb#25805 Fixes scylladb/scylladb#26045	2025-09-18 03:24:43 +02:00
Michał Jadwiszczak	7b9db335c0	db/view/view_building_worker: execute entire `batch::do_work` on tasks shard This change will allow us to get rid of problematic `sharded<abort_source>` and use local `abort_source` instead.	2025-09-18 03:24:43 +02:00
Michał Jadwiszczak	2f65af8aa7	db/view/view_building_worker: store reference to sharded worker in batch Change reference to view building worker in batch to sharded container. In next commits, I'm going to execute `do_work()` exclusively on tasks target shard and sharded reference will be more useful.	2025-09-18 03:24:43 +02:00
Michał Jadwiszczak	e4a0de53ea	db/view/view_building_worker: fix indent	2025-09-18 02:57:36 +02:00
Michał Jadwiszczak	7dfb76f9a7	db/view/view_building_worker: wrap `shared_sstable` in `foreign_ptr` When a staging sstable is registered to view building worker, it needs to make a round trip from its original shard to shard 0 (in order to create a view building task) and back (to be eventually processed). Until now this was done using plain `sstables::shared_sstable` (= `lw_shared_ptr`) which is not safe to be moved between shards. This patch fixes this by wrapping the pointer in `foreign_ptr` and obtains necessary informations (owner shard, last token) on the original shard (instead of on shard0). Then all of those objects are put into freshly introduced structure `staging_sstable_task_info`, which can be safely moved between shards. Fixes scylladb/scylladb#25859	2025-09-18 02:57:36 +02:00
Michał Jadwiszczak	50678030c0	db/view/view_building_worker: use table id in `register_staging_sstable_tasks()` There is no need to pass the pointer only to get id of the table.	2025-09-18 02:57:35 +02:00
Michał Jadwiszczak	b44c223d47	db/view/view_building_worker: move helper functions higher So they can be used in `view_building_worker::register_staging_sstable_tasks()`.	2025-09-18 02:57:35 +02:00
Michał Jadwiszczak	d98237b33c	db/view/view_building_worker: move try-catch outside `invoke_on()` It's just stylist change, to me doing `invoke_on()` in try-catch block looks better than the other way.	2025-09-16 23:15:44 +02:00
Michał Jadwiszczak	9458ceff8f	db/view/view_building_worker: split batch's data preparation and execution The view building batch lives on shard0 but it might be doing work on shard which owns the tablet replica. Until now the batch data was accessed from multiple shards (shard0 and where the batch was executed). This patch fixes this by splitting tasks execution into: - preparation which is always happening on shard0 - actual execution of the tasks on relevant shard, but all necessary data is copied to the shard and batch object isn't accessed. Fixes scylladb/scylladb#25804	2025-09-16 23:13:36 +02:00
Radosław Cybulski	436150eb52	treewide: fix spelling errors Fix spelling errors reported by copilot on github. Remove single use namespace alias. Closes scylladb/scylladb#25960	2025-09-12 15:58:19 +03:00

1 2

57 Commits