mirror of https://github.com/scylladb/scylladb.git synced 2026-06-01 04:26:48 +00:00

Go to file

Kefu Chai b39cc01bb3 compaction_manager: flush all tables before cleanup

according to the document "nodetool cleanup"

> Triggers removal of data that the node no longer owns

currently, scylla performs cleanup by rewriting the sstables. but
commitlog segments may still contain the mutations to the tables
which are dropped during sstable rewriting. when scylla server
restarts, the dirty mutations are replayed to the memtable. if
any of these dirty mutations changes the tables cleaned up. the
stale data are reapplied. this would lead to data resurrection.

so, in this change we following the same model of major compaction:

1. force new active segment,
2. flush all tables
3. perform cleanup using compaction, which rewrites the sstables
   of specified tables

because we already `flush()` all tables in
`cleanup_keyspace_compaction_task_impl::run()`, there is no need to
call `flush()` again, in `table::perform_cleanup_compaction()`, so
the `flush()` call is dropped in this function, and the tests using
this function are updated to call `flush()` manually to preserve
the existing behavior.

there are two callers of `cleanup_keyspace_compaction_task_impl`,

* one is `storage_service::sstable_cleanup_fiber()`, which listens
  for the events fired by topology_state_machine, which is in turn
  driven by, for instance, "/storage_service/cleanup_all" API.
  which cleanup all keyspaces in one after another.
* another is "/storage_service/keyspace_cleanup", which cleans up
  the specified keyspace.

in the first use case, we can force a new active segment for a single
time, so another parameter to the ctor of
`cleanup_keyspace_compaction_task_impl` is introduced to specify if
the `db.flush_all_tables()` call should be skiped.

please note, there are two possible optimizations,

1. force new active segment only if the mutations in it touches the
   tables being cleaned up
2. after forcing new active segment, only flush the (mem)tables
   mutated by the non-active segments

but let's leave them for following-up changes. this change is a
minimal fix for data resurrection issue.

Fixes #16757
Signed-off-by: Kefu Chai <kefu.chai@scylladb.com>

2024-02-01 11:25:53 +08:00

.github

.git: add more skip words

2024-01-29 14:37:03 +02:00

alternator

Merge 'alternator: enable tablets by default if experimental feature is enabled' from Nadav Har'El

2024-01-29 09:22:13 +02:00

api

compaction_manager: flush all tables before cleanup

2024-02-01 11:25:53 +08:00

auth

service/maintenance_mode: move maintenance_socket_enabled definition to seperate file

2024-01-25 15:27:53 +01:00

bin

tools: add cqlsh shortcut

2023-07-12 09:36:59 +03:00

cdc

cdc: not include unused headers

2024-01-11 09:13:37 +02:00

cmake

build: cmake: use # for line comment

2024-01-03 15:05:00 +02:00

compaction

compaction_manager: flush all tables before cleanup

2024-02-01 11:25:53 +08:00

conf

Merge 'Add maintenance socket' from Mikołaj Grzebieluch

2023-12-20 19:04:40 +02:00

cql3

treewide: fix misspellings in code comments

2024-01-31 09:16:10 +02:00

data_dictionary

keyspace_metadata: Drop vector-of-schemas argument from new_keyspace()

2023-12-26 13:00:44 +03:00

Revert "Merge 'Use utils::directories instead of db::config to get dirs' from Patryk Wróbel"

2024-01-31 15:08:14 +03:00

debug

…

dht

db: add formatter for dht::decorated_key and repair_sync_boundary

2024-01-29 11:11:41 +02:00

direct_failure_detector

direct_failure_detector: Avoid throwing exceptions in the success path

2023-03-31 12:40:43 +02:00

dist

scylla_util.py: wait for apt operation on other processes

2023-12-28 19:00:36 +02:00

docs

Merge 'Implement tablet splitting' from Raphael "Raph" Carvalho

2024-01-31 13:59:56 +02:00

exceptions

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

gms

Merge 'Add more logging for gossiper::lock_endpoint and storage_service::handle_state_normal' from Kamil Braun

2024-01-12 10:51:21 +02:00

idl

storage_service: Implement table_load_stats RPC

2024-01-25 18:36:08 -03:00

index

Merge 'scylla-sstable: add support for loading schema of views and indexes' from Botond Dénes

2024-01-24 23:36:54 +02:00

interface

Typos: fix typos in comments

2023-12-02 22:37:22 +02:00

lang

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

licenses

scripts: remove git-archive-all

2023-03-29 18:59:23 +03:00

locator

Merge 'Implement tablet splitting' from Raphael "Raph" Carvalho

2024-01-31 13:59:56 +02:00

message

Merge 'Implement tablet splitting' from Raphael "Raph" Carvalho

2024-01-31 13:59:56 +02:00

mutation

treewide: fix misspellings in code comments

2024-01-31 09:16:10 +02:00

mutation_writer

mutation_writer: do not include unused headers

2024-01-24 15:20:02 +02:00

node_ops

token_metadata: drop the template

2023-12-12 23:19:54 +04:00

raft

Merge 'Add an API to trigger snapshot in Raft servers' from Kamil Braun

2024-01-29 15:06:04 +02:00

readers

reader: do not include unused headers

2024-01-29 16:21:42 +02:00

redis

redis: do not include unused headers

2024-01-31 09:17:18 +02:00

reloc

…

repair

treewide: fix misspellings in code comments

2024-01-31 09:16:10 +02:00

replica

replica: table: pass do_flush to table::perform_cleanup_compaction()

2024-02-01 11:25:53 +08:00

rust

rust: update dependencies

2023-12-17 13:20:25 +02:00

schema

schema: column_mapping::{static,regular}_column_at(): use on_internal_error()

2024-01-31 05:12:33 -05:00

scripts

Typos: fix typos in code

2023-12-13 10:45:21 +02:00

seastar @ 85359b2866

Update seastar submodule

2024-01-22 11:29:50 +01:00

service

compaction_manager: flush all tables before cleanup

2024-02-01 11:25:53 +08:00

sstables

treewide: fix misspellings in code comments

2024-01-31 09:16:10 +02:00

streaming

Merge 'tablets: Add support for removenode and replace handling' from Tomasz Grabiec

2024-01-25 14:49:43 +02:00

swagger-ui @ 12f1da1082

…

tasks

tasks: don't keep internal root tasks after they complete

2024-01-09 13:13:54 +01:00

test

Revert "Merge 'Use utils::directories instead of db::config to get dirs' from Patryk Wróbel"

2024-01-31 15:08:14 +03:00

thrift

thrift: remove unused namespace definition

2024-01-30 09:16:47 +02:00

tools

Revert "Merge 'Use utils::directories instead of db::config to get dirs' from Patryk Wróbel"

2024-01-31 15:08:14 +03:00

tracing

tracing: add formatter for tracing::span_id

2024-01-31 13:43:46 +02:00

transport

Revert "Merge 'Use utils::directories instead of db::config to get dirs' from Patryk Wróbel"

2024-01-31 15:08:14 +03:00

types

utils: do not include unused headers

2024-01-18 12:50:06 +02:00

unified

Update unified/build_unified.sh

2023-12-05 15:23:38 +02:00

utils

Revert "Merge 'Use utils::directories instead of db::config to get dirs' from Patryk Wróbel"

2024-01-31 15:08:14 +03:00

.dockerignore

…

.gitattributes

…

.gitignore

docs: download iam csv files

2023-10-02 12:28:56 +03:00

.gitmodules

Repackaging cqlsh

2023-03-12 20:22:33 +02:00

.gitorderfile

…

.mailmap

…

absl-flat_hash_map.cc

…

absl-flat_hash_map.hh

…

amplify.yml

…

backlog_controller.hh

treewide: apply codespell to the comments in source code

2023-12-20 10:25:03 +02:00

build_mode.hh

release: correct a typo in comment

2023-03-29 13:42:38 +03:00

bytes_ostream.hh

utils/managed_bytes, serializer: add conversion between buffer_view<bytes_ostream> and managed_bytes_view

2023-05-07 17:17:34 +03:00

bytes.cc

bytes: implement formatting helpers using formatter

2023-03-27 20:06:45 +08:00

bytes.hh

bytes.hh: correct spelling of delimiter and delimited

2023-12-18 20:46:21 +02:00

cache_flat_mutation_reader.hh

cache_flat_mutation_reader: fix a broken iterator validity guarantee in ensure_population_lower_bound()

2023-11-16 19:01:18 +01:00

cache_temperature.hh

…

cartesian_product.hh

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

cell_locking.hh

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

checked-file-impl.hh

code: Switch to seastar API level 7

2023-06-06 13:29:16 +03:00

client_data.cc

…

client_data.hh

…

clocks-impl.cc

clocks-impl: format time_point using fmt

2023-11-22 17:44:07 +02:00

clocks-impl.hh

…

clustering_bounds_comparator.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

clustering_interval_set.hh

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

clustering_key_filter.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

clustering_ranges_walker.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

CMakeLists.txt

build: cmake: add "mode_list" target

2023-12-24 12:35:02 +08:00

collection_mutation.cc

Introduce mutation/ module

2023-02-14 11:19:03 +02:00

collection_mutation.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

column_computation.hh

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

combine.hh

…

compound_compat.hh

compound_compat: do not format an sstring with {:d}

2023-07-08 15:13:11 +03:00

compound.hh

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

compress.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

compress.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

concrete_types.hh

make timestamp string format cassandra compatible

2023-07-27 12:01:09 +03:00

configure.py

Revert "Merge 'Use utils::directories instead of db::config to get dirs' from Patryk Wróbel"

2024-01-31 15:08:14 +03:00

CONTRIBUTING.md

Replacing user-group with community forum, added link to U. lesson on Spring Boot Fixed author/email details

2023-02-23 19:05:26 +02:00

converting_mutation_partition_applier.cc

Introduce schema/ module

2023-02-15 11:01:50 +02:00

converting_mutation_partition_applier.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

counters.cc

counters: move fmt::formatter<counter_{shard,cell}_view>::format() to .cc

2023-05-24 09:36:49 +03:00

counters.hh

counters: move fmt::formatter<counter_{shard,cell}_view>::format() to .cc

2023-05-24 09:36:49 +03:00

coverage_excludes.txt

test.py: support code coverage

2024-01-18 11:11:34 +02:00

coverage_sources.list

configure.py support coverage profiles on standrad build modes

2024-01-18 11:11:34 +02:00

cql_serialization_format.hh

…

db_clock.hh

db_clock: specialize fmt::formatter<db_clock::time_point>

2023-04-28 15:48:06 +08:00

debug.cc

…

debug.hh

…

default.nix

build: nix: switch to non-static zstd

2023-02-17 10:29:34 +02:00

Doxyfile

…

duration.cc

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

duration.hh

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

encoding_stats.hh

encoding_state: mark helper methods protected

2023-08-29 15:41:13 +03:00

enum_set.hh

…

fix_system_distributed_tables.py

…

flake.lock

…

flake.nix

…

frozen_schema.cc

Introduce mutation/ module

2023-02-14 11:19:03 +02:00

frozen_schema.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

full_position.hh

Introduce mutation/ module

2023-02-14 11:19:03 +02:00

gc_clock.hh

utils: move hashing related files to utils/ module

2023-02-17 07:19:52 +02:00

gdbinit

…

gen_segmented_compress_params.py

Typos: fix typos in code

2023-12-13 10:45:21 +02:00

generic_server.cc

generic_server: use mutable reference in for_each_gently

2023-11-14 14:25:22 +02:00

generic_server.hh

generic_server: use mutable reference in for_each_gently

2023-11-14 14:25:22 +02:00

HACKING.md

commitlog: use separate directory for schema commitlog

2023-03-30 21:55:50 +04:00

hashing_partition_visitor.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

idl-compiler.py

Typos: fix typos in code

2023-12-13 10:45:21 +02:00

inet_address_vectors.hh

abstract_replication_strategy: calculate_natural_endpoints: make it work with both versions of token_metadata

2023-12-12 23:19:53 +04:00

init.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

init.hh

Merge 'Typos: fix typos in code' from Yaniv Kaul

2023-12-06 07:36:41 +02:00

install-dependencies.sh

build: add crypto++ to dependencies

2024-01-11 16:26:20 +02:00

install.sh

install.sh: use a temporary file when packaging scylla.yaml

2024-01-01 21:50:29 +02:00

interval.hh

interval: make default ctor and make_open_ended_both_sides constexpr

2023-11-06 18:39:53 +01:00

keys.cc

keys: Move exploded_clustering_prefix's operator<< to keys.cc

2023-07-19 11:57:27 +03:00

keys.hh

keys: do not use zip_iterator for printing key components

2023-07-01 23:49:02 +03:00

LICENSE.AGPL

…

log.hh

…

main.cc

Revert "Merge 'Use utils::directories instead of db::config to get dirs' from Patryk Wróbel"

2024-01-31 15:08:14 +03:00

map_difference.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

marshal_exception.hh

…

multishard_mutation_query.cc

reader_permit: store schema_ptr instead of raw schema pointer

2024-01-11 08:37:56 +02:00

multishard_mutation_query.hh

treewide: apply codespell to the comments in source code

2023-12-20 10:25:03 +02:00

mutation_query.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

mutation_query.hh

mutation_query: add formatter for reconcilable_result::printer

2023-11-26 20:20:50 +02:00

noexcept_traits.hh

…

NOTICE.txt

…

ORIGIN

…

partition_builder.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

partition_range_compat.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

partition_slice_builder.cc

partition_slice_builder: add set_specific_ranges()

2023-05-08 07:35:39 -04:00

partition_slice_builder.hh

partition_slice_builder: add set_specific_ranges()

2023-05-08 07:35:39 -04:00

partition_snapshot_reader.hh

Introduce mutation/ module

2023-02-14 11:19:03 +02:00

partition_snapshot_row_cursor.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

protocol_server.hh

…

querier.cc

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

querier.hh

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

query_id.hh

query_id: extract into new header

2023-03-01 10:25:25 +02:00

query_ranges_to_vnodes.cc

everywhere: reduce dependencies on i_partitioner.hh

2023-11-05 20:47:44 +02:00

query_ranges_to_vnodes.hh

everywhere: reduce dependencies on i_partitioner.hh

2023-11-05 20:47:44 +02:00

query_result_merger.hh

…

query-request.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

query-result-reader.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

query-result-set.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

query-result-set.hh

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

query-result-writer.hh

types: move types.{cc,hh} into types

2023-02-19 21:05:45 +02:00

query-result.hh

treewide: do not mark return value const if this has no effect

2023-11-17 17:46:19 +08:00

query.cc

treewide: use #include <seastar/...> for seastar headers

2023-06-06 08:36:09 +03:00

range.hh

…

read_context.hh

compact and remove expired rows from cache on read

2023-06-26 15:29:01 +02:00

reader_concurrency_semaphore.cc

reader_concurrency_semaphore.cc: move stringstream content instead of copying it

2024-01-31 09:31:50 +02:00

reader_concurrency_semaphore.hh

reader_permit: store schema_ptr instead of raw schema pointer

2024-01-11 08:37:56 +02:00

reader_permit.hh

reader_permit: store schema_ptr instead of raw schema pointer

2024-01-11 08:37:56 +02:00

README.md

Replacing user-group with community forum, added link to U. lesson on Spring Boot Fixed author/email details

2023-02-23 19:05:26 +02:00

real_dirty_memory_accounter.hh

real_dirty_memory_accounter: document what the class is doing

2023-05-23 09:11:31 +03:00

release.cc

…

release.hh

…

reversibly_mergeable.hh

…

row_cache.cc

reader: do not include unused headers

2024-01-29 16:21:42 +02:00

row_cache.hh

Merge 'row_cache: abort on exteral_updater::execute errors' from Benny Halevy

2023-10-31 10:07:01 +02:00

schema_mutations.cc

schema_mutations, migration_manager: Ignore empty partitions in per-table digest

2023-07-03 23:06:55 +02:00

schema_mutations.hh

schema_mutations, migration_manager: Ignore empty partitions in per-table digest

2023-07-03 23:06:55 +02:00

schema_upgrader.hh

Introduce mutation/ module

2023-02-14 11:19:03 +02:00

scylla_post_install.sh

dist: drop legacy control group parameters

2023-12-11 19:38:28 +09:00

scylla-gdb.py

reader_permit: store schema_ptr instead of raw schema pointer

2024-01-11 08:37:56 +02:00

SCYLLA-VERSION-GEN

Typos: fix typos in code

2023-12-13 10:45:21 +02:00

seastarx.hh

…

serialization_visitors.hh

…

serializer_impl.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

serializer.cc

utils/managed_bytes, serializer: add conversion between buffer_view<bytes_ostream> and managed_bytes_view

2023-05-07 17:17:34 +03:00

serializer.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

service_permit.hh

…

setup.py

…

shell.nix

…

sstables_loader.cc

sstables_loader: load_new_sstables: auto-enable load-and-stream for tablets

2024-01-16 18:43:52 +02:00

sstables_loader.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

supervisor.hh

…

table_helper.cc

keyspace_metadata: Add default value for new_keyspace's durable_writes

2023-12-26 11:47:37 +03:00

table_helper.hh

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

test.py

test.py: add boost_tests() to suite

2024-01-31 13:43:21 +02:00

timeout_config.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

timeout_config.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

timestamp.hh

…

tombstone_gc_extension.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

tombstone_gc_options.cc

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

tombstone_gc_options.hh

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

tombstone_gc.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

tombstone_gc.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

tox.ini

…

ubsan-suppressions.supp

…

unimplemented.cc

unimplemented: add format_as() for unimplemented::cause

2024-01-19 08:38:30 +02:00

unimplemented.hh

./: not include unused headers

2024-01-17 16:30:14 +02:00

validation.cc

validation: Avoid throwing schema lookup

2023-03-24 08:43:48 +02:00

validation.hh

Introduce schema/ module

2023-02-15 11:01:50 +02:00

version.hh

treewide: use defaulted operator!=() and operator==()

2023-04-27 10:24:46 +03:00

view_info.hh

everywhere: reduce dependencies on i_partitioner.hh

2023-11-05 20:47:44 +02:00

vint-serialization.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

vint-serialization.hh

Typos: fix typos in code

2023-12-05 15:18:11 +02:00

zstd.cc

./: not include unused headers

2024-01-17 16:30:14 +02:00

README.md

Scylla

What is Scylla?

Scylla is the real-time big data database that is API-compatible with Apache Cassandra and Amazon DynamoDB. Scylla embraces a shared-nothing approach that increases throughput and storage capacity to realize order-of-magnitude performance improvements and reduce hardware costs.

For more information, please see the ScyllaDB web site.

Build Prerequisites

Scylla is fairly fussy about its build environment, requiring very recent versions of the C++20 compiler and of many libraries to build. The document HACKING.md includes detailed information on building and developing Scylla, but to get Scylla building quickly on (almost) any build machine, Scylla offers a frozen toolchain, This is a pre-configured Docker image which includes recent versions of all the required compilers, libraries and build tools. Using the frozen toolchain allows you to avoid changing anything in your build machine to meet Scylla's requirements - you just need to meet the frozen toolchain's prerequisites (mostly, Docker or Podman being available).

Building Scylla

Building Scylla with the frozen toolchain dbuild is as easy as:

$ git submodule update --init --force --recursive
$ ./tools/toolchain/dbuild ./configure.py
$ ./tools/toolchain/dbuild ninja build/release/scylla

For further information, please see:

Developer documentation for more information on building Scylla.
Build documentation on how to build Scylla binaries, tests, and packages.
Docker image build documentation for information on how to build Docker images.

Running Scylla

To start Scylla server, run:

$ ./tools/toolchain/dbuild ./build/release/scylla --workdir tmp --smp 1 --developer-mode 1

This will start a Scylla node with one CPU core allocated to it and data files stored in the tmp directory. The --developer-mode is needed to disable the various checks Scylla performs at startup to ensure the machine is configured for maximum performance (not relevant on development workstations). Please note that you need to run Scylla with dbuild if you built it with the frozen toolchain.

For more run options, run:

$ ./tools/toolchain/dbuild ./build/release/scylla --help

Testing

See test.py manual.

Scylla APIs and compatibility

By default, Scylla is compatible with Apache Cassandra and its APIs - CQL and Thrift. There is also support for the API of Amazon DynamoDB™, which needs to be enabled and configured in order to be used. For more information on how to enable the DynamoDB™ API in Scylla, and the current compatibility of this feature as well as Scylla-specific extensions, see Alternator and Getting started with Alternator.

Documentation

Documentation can be found here. Seastar documentation can be found here. User documentation can be found here.

Training

Training material and online courses can be found at Scylla University. The courses are free, self-paced and include hands-on examples. They cover a variety of topics including Scylla data modeling, administration, architecture, basic NoSQL concepts, using drivers for application development, Scylla setup, failover, compactions, multi-datacenters and how Scylla integrates with third-party applications.

Contributing to Scylla

If you want to report a bug or submit a pull request or a patch, please read the contribution guidelines.

If you are a developer working on Scylla, please read the developer guidelines.

Contact

The community forum and Slack channel are for users to discuss configuration, management, and operations of the ScyllaDB open source.
The developers mailing list is for developers and people interested in following the development of ScyllaDB to discuss technical topics.

Languages

C++ 72.3%

Python 26.5%

CMake 0.3%

GAP 0.3%

Shell 0.3%